Utility of Tablet-Based Eye Tracking for Early Screening of Poststroke Cognitive Impairment: Diagnostic Cohort Study.
Authors: Xie B, Ren R, Zhang Y, Wang X, Zhen J
Journal: JMIR mHealth and uHealth
mental health
psychology
open access
Abstract
Recent advances in large language models (LLMs) have demonstrated their transformative potential across diverse health care domains. These applications range from automating medical documentation and patient education to providing personalized health consultation and drug discovery [-]. Despite this broad use, AI-driven diagnostic assistance remains one of the most promising and impactful scenarios in clinical practice [-]. Given the critical and complex nature of clinical decision-making, exceptionally high accuracy is required. However, diagnostic reasoning in real-world practice often involves uncertainty, incomplete information, and iterative hypothesis refinement []. In this context, the application of LLMs raises concerns not only about “hallucinations” but also about their reliability and stability in supporting complex clinical reasoning processes []. Therefore, their diagnostic performance must be carefully and systematically evaluated before being considered for clinical use. China’s LLM development is progressing rapidly with strong national support. However, most existing studies focus on evaluating LLMs’ diagnostic capabilities in English [-]. Given that some research indicates ChatGPT performs better with English input compared to Chinese, and that Chinese LLMs leverage large Chinese corpora and training data representative of the Chinese population, evaluating their diagnostic ability in the Chinese context is essential for their application in Chinese clinical settings. This necessity is further underscored by the unique challenges of Chinese medical natural language processing (NLP), where models must navigate highly unstructured clinical notes characterized by complex syntactic structures and nonstandardized medical abbreviations. Although some studies have explored Chinese LLMs’ diagnostic ability in Chinese contexts, research in this area remains limited []. Early assessments of LLMs’ diagnostic competence often used multiple choice questions (MCQs), which may not accurately reflect real-world clinical performance due to their structured nature. Recent studies have shifted toward clinical vignette-based evaluations that better mimic actual clinical scenarios [-]. Most of these evaluations provide LLMs with complete patient data simultaneously. However, real clinical reasoning typically follows a hypothetico-deductive process, where clinicians generate an initial differential diagnosis (DDx) from limited information, then iteratively refine these hypotheses with additional data. Assessment methods simulating this incremental information provision have demonstrated reduced diagnostic accuracy in LLMs compared to approaches presenting all data at once.