Dwarf mongooses pre-emptively alter their behaviour relative to the threat posed by different rival groups.
Authors: Arbon JJ, Morris-Drake A, Kern JM, Radford AN
Journal: Nature ecology & evolution
mental health
psychology
open access
Abstract
Medical artificial intelligence (AI) has immense potential to improve health outcomes, particularly in regions in which specialized medical expertise is scarce. At the same time, AI also poses new challenges and risks, including security vulnerabilities that arise when models are deployed. Untrusted users with access to an AI model may, by merely observing its predictions, steal its parameters or perform privacy attacks, which can extract sensitive details about the data used for model training. Privacy attacks against an AI model can enable detailed inferences about the individuals who contributed to its training data. For example, a membership inference attack (MIA) attempts to determine whether the data of a specific patient were included in the training dataset of a model. The extent to which this constitutes a privacy violation is nuanced and depends on factors such as the underlying training population and the deployment context of the model. Although inferring membership for a model trained on a general population may be benign, doing so for a model trained on a narrow, disease- or centre-specific cohort acts as a direct proxy for sensitive medical information. For example, a successful MIA against the model in ref. , which predicts anti-cancer immunotherapy efficacy from routine blood test data, reveals that an individual has cancer. The accelerating deployment of medical AI models trained on sensitive patient data calls for rigorous privacy risk assessments. However, previous studies primarily quantified the success rate of MIAs, in aggregate, across all records in a training dataset. This implicitly averages risk across records, thereby obscuring important information on record- and patient-level attack success. Consequently, the risk that an individual faces by contributing their personal data (often multiple records) to an AI training dataset is poorly understood. Given that medical data are a key target for cybercriminals, and pseudonymization alone is increasingly recognized as insufficient to prevent the re-identification of individuals in large, high-dimensional datasets, there is a need to improve our understanding of the threat that AI privacy attacks pose to individual patients.