Digital Media Literacy Group for Adolescents in Child and Adolescent Psychiatry: Prospective Multicenter Single-Arm Feasibility Study.
Authors: Petras IK, Wüllner S, Hermenau K, Meerkamp K, Siniatchkin M
Journal: JMIR formative research
mental health
psychology
open access
Abstract
The integration of massive, multiparty patient data is crucial for developing high-quality medical AI. Compared to single-site datasets, those aggregated from multiple institutions exhibit greater diversity in patient demographics, clinical practices, and health care conditions. This enables researchers to obtain sufficient samples to train robust models for specific tasks, such as studying rare diseases []; achieve results with higher reproducibility and generalizability []; accelerate the translation of research findings into clinical practice []; and discover novel patterns that are undetectable in any single dataset []. As a result, such integration significantly promotes the progress of medical data analysis and makes their applications widespread in many real-world clinical scenarios. However, the sharing of such data is severely constrained by various privacy regulations, such as the HIPAA (Health Insurance Portability and Accountability Act) Privacy Rule in the United States and the GDPR (General Data Protection Regulation) in the European Union [,]. These regulations, coupled with the severe repercussions of privacy breaches—including reputational damage, substantial financial penalties, and health care fraud—make medical institutions highly reluctant to share their sensitive patient data []. Consequently, there is an urgent need for secure multiparty medical AI analysis protocols that can navigate the nontrivial trade-off between practical utility and robust security. Most existing privacy-preserving AI protocols are designed for, or evaluated on, cross-sectional data. This paper, in contrast, addresses the challenge of analyzing longitudinal patient data using recurrent neural networks (RNNs) in a privacy-preserving manner. Longitudinal patient data, which records patient visits, test results, and treatments over time, provides more invaluable insights into the relationship between temporal patterns of features and events of interest []. RNNs are particularly well-suited for mining such data due to their ability to model nonlinear relationships and high-dimensional temporal dependencies. They have proven essential in advancing personalized medicine through disease progression modeling [], outcome prediction [], and treatment recommendation []. Despite their value, longitudinal datasets are often harder to collect and tend to be smaller in scale than their cross-sectional counterparts. This scarcity frequently necessitates pooling data from multiple providers to train robust models, a process thwarted by the distributed and sensitive nature of the data itself.