← Back to Research Papers

Measuring Depression Severity With Clinical Global Impression-Severity Scale Scores From Clinical Notes Using Large Language Models: Validation Study.

Authors: Li K, Zirikly A, Collica SC, Goes FS, Zhao C, Nguyen T, Gagliardi JP, Goldstein BA, Hong H, Stuart EA, Zandi PP
Journal: JMIR formative research
mental health psychology open access

Abstract

AI technologies are increasingly being integrated into health care systems and are transforming the delivery of medical services []. Evolving from early rule-based systems to modern approaches such as machine learning, deep learning, and natural language processing, AI can analyze large data sets to identify patterns and relationships beyond human perception, with recent advances such as large language models further extending its capabilities []. Their applications have expanded beyond early use in diagnostic support to a wide range of areas, including disease screening, treatment decisions, chronic disease management, medication consultation, and remote health monitoring []. AI-enabled technologies are developing rapidly and are becoming more embedded in health care decision-making and service delivery. Although their clinical use remained limited in 2021, with only a small proportion of AI technologies implemented in practice [], adoption has expanded rapidly in recent years. By March 2024, approximately 79% of health care organizations reported using AI technologies in the Microsoft-International Data Corporation study []. At the same time, growing evidence indicates that AI-based approaches can achieve performance and convenience comparable to, or even exceeding, those of traditional methods [,]. As AI increasingly permeates health care, a range of concerns has emerged. Stakeholders worry that AI systems cannot fully replicate clinician decision-making, as they may overlook patient cognitive status, quality of life, and individual preferences []. Additional concerns include the lack of transparency and explainability of algorithms, potential errors or malfunctions, overreliance on technology that could dehumanize patient care [], and data privacy and security issues and other potential challenges []. These multidimensional characteristics may influence trust, acceptance, and ultimately the adoption of AI-enabled health care services. Understanding how these characteristics are perceived and which attributes are most valued by patients, professionals, and other stakeholders is essential for promoting the adoption of AI in health care and ensuring that AI-enabled technologies are better aligned with real-world needs, especially for patients and implementation contexts. Discrete choice experiments (DCEs), grounded in random utility theory, are widely used to elicit preferences by asking individuals to choose between hypothetical alternatives that vary across multiple attributes and levels [,]. By requiring individuals to select between alternatives, DCEs allow researchers to quantify trade-offs between service characteristics and have been shown to approximate real-world decision-making, correctly predicting more than 93% of choices []. Given the multidimensional nature of AI-enabled health care technologies, this approach is particularly useful for identifying the attributes that stakeholders value most when evaluating such innovations. Beyond preference estimation, the validity and comparability of DCE evidence depend on transparent reporting of methodological procedures. Incomplete reporting of key elements, including attribute development, experimental design, randomization, and econometric modeling, may limit the credibility, reproducibility, and synthesis of findings [,]. This issue is especially relevant in AI-enabled health care, where the “black-box” nature of AI systems has heightened concerns regarding transparency, trustworthiness, and real-world implementation []. Transparent reporting is therefore essential for interpreting preference evidence and supporting evidence synthesis.