Lavandula angustifolia and Echium amoenum as emerging adjuncts to cognitive behavioral therapy for major depressive disorder: mechanistic insights into neurotrophic, monoaminergic, and anti-inflammato
Authors: Song Y, Yang J, Xiao F
Journal: Frontiers in psychiatry
mental health
psychology
open access
Abstract
The emergence of multimodal artificial intelligence (MMAI) offers many reasons for excitement, representing more than an incremental advance in predictive capacity. Most AI systems have been historically limited to unimodal analysis, interpreting one data type at a time, and therefore struggle to capture broader contextual relationships. This is quickly changing as new systems leverage data fusion to embed multiple data streams within a shared representational space, enabling richer, cross-modal reasoning. In healthcare, MMAI is being used to study health and disease in context, situating findings among behavioral, environmental, and physiological contingencies, illuminating underlying mechanisms such as neural or genetic contributors to disease, and promising gains in diagnostic accuracy, early detection, and more holistic patient support [, ]. When applied in medicine, MMAI holds the potential to transform how biomedical knowledge and data products are produced. It does this by dissolving a foundational distinction between data that records clinically relevant signals from a patient (e.g., physiological or psychometric readings, tissue samples, a clinician's direct observations) and data that are produced about a patient by the computational systems and the infrastructures in which they are embedded. Existing frameworks for privacy, consent, and accountability were built on the assumption that health data are of the first kind, i.e., observed, measured, or disclosed, with governance focused on controlling who can access and use those data. MMAI increasingly produces data of the second kind, i.e., information that did not exist independent of an inferential process yet enters clinical and research records with a similar epistemological standing as observed data. MMAI is ethically novel not only because it infers across modalities, but because it can use generative AI to convert those inferences into new clinical data objects that enter records and workflows with the practical status of observed fact. When these new data objects enter the clinical record unaccompanied by reliable markers of their inferential origin, they become indistinguishable from observed facts. Unlike earlier forms of synthetic data, which are generated intentionally and usually labeled as such in the context of their production, MMAI-generated data objects are produced as part of clinical or analytical workflows where their inferential origin might not be tracked, labeled, or preserved. For purposes of this paper, an “MMAI-generated data object” is a model-produced or model-mediated output derived from heterogeneous data streams that becomes sufficiently durable or operational to be stored, displayed, transmitted, incorporated into a record or dataset, reused in later analysis, or used to guide clinical interpretation or intervention. The category includes synthetic images, AI-generated clinical text, inferred scores or attributes, latent embeddings, and hybrid outputs that combine model generation with human review, but excludes every transient internal model state or ordinary probabilistic inference. Generative AI refers here to computational systems that produce new outputs, such as text, images, scores, synthetic features, or latent representations, rather than only classifying or retrieving existing data []. These outputs can be produced by several model families, including generative adversarial networks, diffusion models, transformer-based language models, and multimodal or foundation models [–]. Multimodal AI refers to systems that integrate, align, or translate across multiple data modalities, such as imaging, text, audio, sensor streams, physiology, or structured clinical data [, ]. The two categories overlap but are not identical; a system can be generative without being multimodal, multimodal without producing durable generated outputs, or both generative and multimodal. The ethical concern arises when these outputs circulate without clear information about how they were generated, which modalities shaped them, how uncertain they are, and whether they have been validated for later clinical or research use.