← Back to Research Papers

Sperm-Associated Mercury Species Levels Determine Its Impairment: Evidence from Paired Analyses of Whole Semen and Seminal Plasma.

Authors: Yuan M, Tsui MT, Xu M, Wang CC, Chan DYL
Journal: Biological trace element research
mental health psychology open access

Abstract

Large language models (LLMs) are becoming widely used in applications ranging from conversational agents to decision-making systems capable of consequential decisions, such as providing medical advice [], legal advice [] and mental well-being support []. In their daily lives, early adopters are already utilizing chatbots to inform their personal decision-making []. Moreover, a substantial proportion of individuals appear willing to delegate moral decisions to AIs []. As humans increasingly interact with LLMs, understanding our ability to detect and align with LLMs’ judgments becomes crucial, particularly given the risk of misuse, such as the dissemination of disinformation by LLM-powered bots [,]. Determining whether a decision is made by a human or a machine is crucial: it enhances safety by revealing our susceptibility to manipulation, possesses intrinsic epistemic value [], and guides the development of LLMs towards human- preferred outputs []. Concerningly, recent studies indicate that humans often struggle to reliably distinguish AI-generated texts from human texts, across diverse contexts such as poetry [–] and media misinformation []. Additionally, strategies can be employed to “humanize” AI-generated content to increase the difficulty of detection. “Humanized” LLMs have sometimes been judged as more human-like than actual human-generated responses []. Before LLMs, research into applied fields such as autonomous vehicles highlighted the need for alignment in human and machine moral decision making []. Prior research has shown a human tendency to favor human-generated decisions over machine-generated ones, a phenomenon known as algorithm aversion []. However, this phenomenon is context-dependent []. For instance, humans tend to prefer human judgment over AI judgment in the context of medical decision-making [], but prefer AI judgment in numerical tasks, such as song rank forecasts based on big data methods [].