← Back to Research Papers

Best Evidence Summary for Improving Self-Awareness of Falls in Elderly Postoperative Patients with Osteoporotic Vertebral Compression Fracture.

Authors: Huo Y, Wang X, Wang Y, Yang J, Zeng Y, Du H
Journal: Journal of multidisciplinary healthcare
mental health psychology open access

Abstract

Digital twins are increasingly proposed as tools for personalized medicine, with the aim of creating computational representations of patients that integrate data, update over time, and support prognosis, monitoring, simulation, or therapeutic planning. In medicine, however, the term “digital twin” covers heterogeneous systems, including descriptive replicas, synchronized simulators, mechanistic models, machine-learning forecasters, in silico trial platforms, and clinical decision-support tools. This heterogeneity is not a weakness: different clinical uses require different architectures and validation standards. The central issue is therefore not to impose a universal definition of medical digital twins, but to clarify what evidence is required for a given use case. Predictive performance remains necessary. A digital twin that cannot reproduce observed trajectories, remain calibrated, or anticipate clinically relevant outcomes is unlikely to be useful. However, predictive accuracy evaluated under historical clinical practice may be insufficient when the twin is used to compare interventions, treatment timings, dosages, or sequential care strategies. In such settings, the relevant question is not only what is likely to happen under current practice, but what is predicted to happen under specified alternative actions, assumptions, and time horizons. This distinction is well recognized across causal inference, uncertainty quantification, decision theory, model validation, dynamic treatment regimes, control theory, and forecast verification. The contribution of this article is therefore not to introduce a new validation theory or to claim that all digital twins must be causal models or scientific instruments. Rather, it synthesizes established validation principles for a specific subset of medical digital twins: systems dynamically linked to patients or populations, updated with data over time, and used to support intervention-oriented clinical decision-making. We use the term “scientific instrument” pragmatically, as a validation lens rather than as a new definition of digital twins. When a digital twin informs clinical action, its outputs should be interpreted as conditional, model-mediated evidence whose reliability depends on assumptions, calibration, operating range, uncertainty representation, updating stability, and context of use. This framing does not imply that all medical digital twins require the same evidentiary burden. Twins used for visualization, education, descriptive monitoring, or short-term forecasting may be evaluated primarily through predictive performance, calibration, usability, and robustness. Twins used to recommend, compare, or optimize clinical actions require additional evidence because their outputs may influence treatment decisions and patient trajectories. Recent work has explicitly framed digital twins as part of the broader P4 medicine paradigm, emphasizing their role in predictive, preventive, personalized, and participatory care. In this view, digital twins are not a single method but a flexible paradigm combining patient-specific data, mechanistic modeling, dynamic updating, virtual intervention, explainability, and learning over time. This article is consistent with this broad framing but focuses on a narrower question: what validation evidence is required when such systems are used to support intervention-oriented clinical decisions? The central principle is that the validation burden should match the decision claim. For intervention-oriented digital twins, evaluation should extend beyond scalar accuracy metrics to include verification, validation for the intended use, uncertainty quantification, robustness under regime change, stability of updating, and decision-level consequences. The relevant question is therefore not simply whether the model is accurate, but for what decision, under what assumptions, over what time horizon, and within what operating range the twin can be considered valid.