← Back to Research Papers

Intellectual Humility and Probabilistic Belief.

Authors: Nelson L, McGregor K, Lombrozo T
Journal: Cognitive science
mental health psychology open access

Abstract

Understanding the clinical course of psychiatric disorders is essential for effective clinical decision-making and treatment planning [-]. Clinical course information includes key temporal and event-based details, such as the onset date, the number of past episodes or relapses, and the frequency of psychiatric hospitalizations. Extracting this information from free-text notes is challenging due to the various ways it is documented (eg, “first psychotic break at 19,” “multiple depressive episodes since 2015,” and no prior psychiatric hospitalizations). Manual retrieval from medical charts is both time-consuming and labor-intensive. Natural language processing (NLP) offers a solution by automating the extraction of these critical details, thereby improving efficiency and reducing human error []. Early NLP approaches combined rule-based algorithms with machine learning. For instance, 1 study developed a hybrid NLP system that extracted time expressions and classified relevant text to generate a ranked timeline of probable psychosis onset dates from mental health records []. Clinical course information in psychiatric discharge summaries is often expressed in unstructured free text with substantial linguistic variability. For example, the same clinical concept may be documented as “first psychotic break at age 19,” “onset in 2009,” or “symptoms dating back to his early twenties.” Rule-based NLP systems require manually specified patterns for different expression types and may not generalize well to heterogeneous descriptions of psychiatric symptoms, episodes, hospitalizations, and temporal relationships. In addition, although the discharge summaries in this study were predominantly written in English, some records contained Chinese expressions, transliterated terms, or code-mixed content, which further complicated purely rule-based extraction. Large language models (LLMs) are well-suited to this task because they can interpret variable free-text expressions via instruction-following and can be adapted to domain-specific annotation schemas through fine-tuning, as demonstrated by the extraction performance observed in this study. Recent studies demonstrate that LLMs can effectively parse clinical text and answer such questions with reasonable accuracy [-]. One study used zero-shot learning with Flan-T5 to extract specifiers, such as severity and remission, for substance use disorders, outperforming traditional rule-based methods in recall by capturing diverse linguistic variations []. This study aimed to evaluate locally deployable, open-source LLMs for extracting clinical course information from psychiatric discharge summaries. We developed a 2-stage framework that first extracts sentence-level clinical events and temporal information and then predicts 4 chart-level features: first episode onset time, episode count, number of psychiatric hospitalizations, and most recent hospitalization time. We compared this framework with direct and joint extraction approaches to assess whether explicit sentence-level extraction improves chart-level prediction.