Healthcare interventions to improve health outcomes for racially/ethnic minoritised people with multiple long-term conditions: a systematic review and narrative synthesis.
Authors: Hayanga B, Joshi M, Hartley K, Oyibo P, Sarwar R, Tariq S, Kashyap M, Bécares L
Journal: BMC public health
mental health
psychology
open access
Abstract
Longitudinal clinical cohort data are invaluable for tracking the evolving needs and outcomes of individuals with chronic conditions. For people living with HIV (PWH), long-term clinical cohorts have been central to major advances in treatment development and epidemiological surveillance. Continued progress in these areas hinges on the ability to share rich, individual-level clinical trajectories collected from diverse global sources. However, stringent and incongruent privacy regulations, ranging from the European Union’s General Data Protection Regulation (GDPR) to Brazil’s Lei Geral de Protecao de Dados, have erected a patchwork of firewalls that hinder data sharing. Even the most successful international research consortia approve only a handful of narrowly scoped data requests each year, slowing knowledge discovery and dissemination, and ultimately delaying clinical benefits. Synthetic health data have emerged as a potentially promising way to enable data sharing and collaboration across countries. When generated properly, synthetic data can mimic the statistical properties of source databases without linking produced records to real people, offering a pragmatic path to easier data sharing that preserves privacy. Generative AI technologies have made it possible to produce realistic cross-sectional snapshots of electronic health record (EHR) data and simple longitudinal sequences, enabling critical applications, such as trial-based association studies, outcome predictive modeling, geotemporal epidemic forecasting, and DNA-disease association analysis—while researchers navigate complex administrative processes required to access real data. However, generating high-quality synthetic longitudinal clinical cohort data faces non-trivial challenges. Cohorts of people with chronic conditions like HIV typically span decades and include critical time-to-event outcomes, such as a variety of clinical endpoints (e.g., diagnoses of comorbidities), medication changes, treatment failures, and mortality. These complexities are fundamental to ensuring clinical relevance and usability of longitudinal data. Yet current synthetic data generation methods struggle to capture the temporal depth and clinical complexity of chronic condition trajectories. Most of these methods are developed and evaluated with short horizons (e.g., 24-h ICU streams or datasets with a limited number of visits), which can reduce a marathon-length chronic condition journey to a snapshot. Another challenge arises from the need to process and generate a mix of continuous (e.g., CD4 cell count measurements), discrete (e.g., test frequencies), and categorical (e.g., diagnostic codes) variables, while simultaneously preserving their temporal relationships. Existing methods, mostly adapted from unimodal domains like imaging or natural language processing, rely on tokenizing continuous and discrete values into coarse categories. This process potentially obscures clinically meaningful variations. Additionally, current synthetic data evaluation strategies typically rely on generic resemblance metrics. These are limited in scope. In particular, they do not assess whether scientifically relevant relationships are preserved in synthetic cohort data, which would be a much more demanding standard of data quality. For instance, an essential but missing metric is whether models fitted on synthetic data can reproduce time-to-event patterns observed in real cohorts, an outcome that reflects the analytical value of longitudinal clinical cohort data and is essential for evaluating treatment effects and informing health policy. Another overlooked metric examines whether longitudinal synthetic data can yield risk factor estimates comparable to those derived from real data, which addresses the capability of synthetic data to support scientific hypothesis generation.