GLP-1 receptor agonists and psychiatric outcomes: a bibliometric analysis of research trends, thematic evolution, and collaboration networks.
Authors: Roca Mora MM, Sanchez-Monge M, Balardini Ayoub D, Aguiar-Barros ABP, Serrano-Porras S, de Filippis R, Schoretsanitis G
Journal: Frontiers in psychiatry
mental health
psychology
open access
Abstract
In neuroscience, behavior remains the primary outcome of interest across many diverse subfields. As such, quantitative and unbiased assessments of animal behavior are critical, yet traditional scoring methods are not only laborious, but also prone to human errors and biased by subjective interpretation, even with corrective efforts. Over the past decade, computer vision methods—most notably convolutional neural network (CNN)-based frameworks such as DeepLabCut, SLEAP, and Lighting Pose—have revolutionized behavioral studies by enabling markerless pose estimation in 2D and 3D. When combined with clustering algorithms that group keypoints into behaviorally meaningful states, these methods have enabled supervised and unsupervised classification of specific behaviors. These approaches represent an important step toward standardization and reproducibility in behavioral neuroscience. However, pose-based clustering alone remains insufficient for understanding how animals interact with the environment—particularly in trial-based tasks where outcomes (e.g., hits versus misses) are critical. Tools such as capacitive sensors, IR beam breaks, and motor encoders are well suited for low-dimensional tasks (e.g., lever pressing) but introduce significant constraints in less restrictive paradigms. Rotary encoders require a physical manipulandum, which can alter natural reach kinematics and may obstruct the visual field. Similarly, lick sensors and infrared beam-breaks primarily measure physical contact or movement rather than task success; these binary signals provide only an indirect proxy that cannot reliably report functional trial outcomes. One such task is the skilled forelimb reach-to-grasp and retrieval assay; in this context, true task success cannot be reliably inferred from simple binary signals (water spout contact), which fail to capture the full, integrated sequence of reaching, retrieval, and consumption. For example, a lick sensor may register a "false positive" if the animal contacts the spout without successfully retrieving the water, or a "false negative" if the animal retrieves and consumes the water directly from its paw without triggering the sensor. Consequently, sensor-based approaches may not be suitable for all behavioral paradigms. At present, there are no robust out-of-the-box video-based approaches that do not require fine-tuning that can distinguish successful from failed reaching attempts. By contrast, video-based LLM scoring evaluates the full, goal-directed behavioral sequence, enabling accurate outcome classification without imposing physical constraints or relying on indirect, approximated signals. In recent years, the rapid evolution of large multimodal models (LMMs) has opened new frontiers in video understanding. Many LLM-based models, such as VideoPrism from Google DeepMind and MouseGPT, can achieve human-level performance in pose estimation and self-directed behavior classification from raw videos alone, reaching parity with expert scorers on benchmark datasets of mouse videos. These advances suggest that general-purpose video LLMs trained on large and diverse datasets may go beyond recognizing self-directed behaviors to classifying environmental interactions and their resulting outcomes. However, these approaches typically require extensive fine-tuning of downstream video encoders or task-specific classifiers to generalize across different views or behavioral paradigms. Most recently, the release of Gemini 2.5 Pro and Qwen3-VL models surpasses prior models across a range of video understanding benchmarks and raises the possibility that video LLMs could generalize to mouse behavior scoring without additional training. Here, we introduce a workflow that leverages video LLMs to directly segment and score rodent reach-to-grasp behaviors from single-view videos. This approach lays the foundation for providing an efficient and scalable alternative to traditional pipelines that rely on pose estimation, feature engineering, and clustering. Using this framework, we benchmark several recent video LLMs on head-fixed mice performing a water-reaching task. Our results suggest that video LLMs may streamline behavior analysis, reduce hardware and annotation overhead, and offer fully generalizable scoring capabilities for behavioral neuroscience.