Clinical indicators associated with pain trajectories in non-specific neck pain: a secondary analysis of a multicentre prospective observational cohort to develop a clinical compass for personalised p
Authors: Verwoerd MJ, Felius RAW, Konings S, Kiers H, Smeets RJEM
Journal: BMJ open
mental health
psychology
open access
Abstract
The management of unruptured intracranial aneurysms (UIAs) is among the most finely balanced decisions in neurovascular practice, requiring the aneurysmal rupture risk to be weighed against the procedural hazards of treatment, a judgement most often entrusted to neurovascular multidisciplinary teams (MDTs) []. The context in which patients meet these decisions, however, has changed: patients increasingly consult large language models (LLMs) such as ChatGPT, Gemini and Claude to interpret their diagnosis and form a view on treatment before reaching specialist review. Importantly, the depersonalised, structured prompts through which clinicians typically evaluate these models bear little resemblance to the way patients actually query them, and clinicians’ impressions of how LLMs perform may therefore not reflect the advice their patients receive. Patients do not query these models in a single, impersonal exchange; instead, they present in the first person, across unstructured and often multi-turn conversations [], embedded with idiosyncrasies and characterised by their own emotional stance which can include anxiety about undergoing treatment or about leaving the aneurysm untreated. Existing literature has benchmarked LLM recommendations against expert judgement in UIA care, consistently finding no better than fair-to-moderate agreement alongside recognisable, model-specific treatment tendencies []. However, such studies repeatedly use third-person prompting structures which are inconsistent with real patient use. The lack of primary literature on the effects of more practical personalised prompting represents a critical gap, as LLMs can exhibit sycophantic tendencies tied to their design of being optimised through user-satisfaction-driven reinforcement learning [–]. As a result, there is a tendency of LLMs to tell users what they appear to want to hear, and query phrasing alone can drive a model to opposite conclusions on the same medical question. Whether this framing sensitivity extends to a genuinely equipoise-laden management decision, and, if it does, whether a patient’s persona and emotional state pull the recommendation toward or away from expert consensus, remains unknown. We therefore conducted a comparative benchmarking study to quantify how patient persona and affective framing alter the UIA management recommendations of three frontier LLMs. Primarily, we compared, head-to-head, the recommendations of ChatGPT, Gemini and Claude for 67 MDT-adjudicated UIA cases presented under four prompting conditions including a third-person clinical vignette, a neutral first-person account, and first-person accounts anxious about treatment and about non-treatment respectively, each benchmarked against neurovascular MDT consensus. Secondarily, we analysed the lexical choices and phraseology employed by LLMs under anxiety framings, and we characterised the direction of any framing-induced shifts, relative to expert consensus, to ascertain potential clinical risks to patients who could be carrying AI-shaped expectations.