Large-scale analysis demonstrates the influence of CYP2C19 genotype on specific SSRI side effects.
Authors: Eijsbouts C, Jiang Y, Ashenhurst JR, Granka JM, 23andMe Research Team, Pitts S, Auton A, Abul-Husn NS, Chubb A, Wu RR
Journal: The pharmacogenomics journal
mental health
psychology
open access
Abstract
The use of generative AI has become ubiquitous. As individuals increasingly incorporate AI into daily life – in search engines and tools for everyday correspondence – researchers face new challenges for both collecting and validating data in online survey research. As AI services evolve, the models become better informed and coherent; AI-powered search, chat, and automated tasking can easily be used to help respond to online questionnaires even with a genuine human at the helm. As researchers who work with human research participants, we are concerned about the serious implications of this for the reliability of data. AI technologies provide individuals with tools that expand human capability, but their widespread use can challenge researchers’ ability to capture true public understanding and opinions on critical scientific and social topics. The problem of agentic “bots” completing surveys is well-described; and a recent article in has described that such chatbots may already be indistinguishable from actual human participants. Adding to the researcher’s list of woes is our concern here – that as easy and prolific access to these tools increases, genuine survey respondents (i.e., actual human participants engaging with a survey) are also using AI to generate some of their responses. The increasingly accessible embedding of AI assistant tools into word processing software, smartphones, and browser software makes it increasingly easy for survey respondents to access these tools while responding to a survey. Such practices will inevitably lead to research that aligns more closely with search engine outputs and language model predictions than with diverse perspectives, societal discourse, and public opinion. Indeed, AI-generated responses to surveys may not even be a deliberate form of “cheating” by respondents – rather, they may believe that AI is a means of generating “correct” responses to questions they are asked. This human-guided, intentional use of AI to “improve” survey responses poses another significant challenge for researchers seeking to improve the validity of their survey responses. There has been considerable methodological attention paid to potential use cases for AI to assist in the gathering of data for research. Yet, as AI becomes increasingly integrated into data collection, it may give rise to a dystopian scenario in which social science becomes an echo chamber of AI-generated sources—where AI both asks and answers questions, and perhaps later reviews the paper—resulting in insights and conclusions recycled from programmed responses, even nonsensical responses, rather than reflecting authentic human experiences and perspectives. If left unchecked, entire datasets could essentially consist of imputed variables, completely reliant on artificial data rather than genuine human input. This growing trend not only threatens the validity of online survey data but also the invaluable insights used to inform policy, programs, and communications.