Temporal Dynamics of Internal vs External Entrapment in Suicidal Ideation in Psychotherapy Outpatients: Prospective Longitudinal Cohort Study Using Ecological Momentary Assessment.
Authors: Wichelhaus E, Schreiber D, Höller I, Forkmann T
Journal: Journal of medical Internet research
mental health
psychology
open access
Abstract
In text classification tasks, researchers require a , either written from scratch or adapted for a new setting from an existing one (For example, the promise classification codebook developed by [] for a trust game was later adapted by [] for a public goods game.) to instruct annotators. These codebooks lay out the underlying context and theory for the classifications, and provide detailed category definitions and annotated examples for each category. When researchers turn to large language models (LLMs) to perform the same task, a natural question arises: can the codebook prepared for human annotators be used directly as an LLM prompt? The literature offers guidance on how to prompt “effectively”. Researchers are advised to structure their prompts with markdown headers and bullet lists [–], to assign role personas [–], to provide background context and label descriptions [,], to include worked examples [], and to restructure the codebook for the model []. However the majority of this advice draws on early prompting research on GPT-2 and GPT-3 [,], where models were less capable at following instructions and prompts were typically short, one- or two-sentence task descriptions rather than detailed codebooks []. Following these recommendations requires effort to transform the codebook into something “optimised” for the model, yet whether they still apply to today’s frontier models, or to substantially more detailed prompts, is unclear. Apart from the informational content of a prompt, there is also the question of how that content is presented: choices about formatting (markdown vs. plain text, bullet lists vs. prose), framing (role persona, task preamble), and phrasing (specific word choices, sentence structure). We refer to these collectively as : changes that alter how the content is presented without changing the informational detail of what it communicates about the classification task. Ideally, classification outcomes should be robust to surface variations. Yet, several studies have shown that LLM outputs can be sensitive to them, producing significant variance in classification outcomes [,]. We refer to this sensitivity as : the dependence of classification outcomes on surface variations []. If models are brittle in this way, then any particular prompt formulation may not be robust, and iterating on phrasing and format may not lead to a reliably optimal prompting solution.