← Back to Research Papers

Communicating Diversity: (Cognitive) Ableism in Information-Seeking Research.

Authors: Cowell RL
Journal: Autism in adulthood
mental health psychology open access

Abstract

Recent advances in the field of natural language processing have showcased the exceptional performance of large language models (LLMs) across various natural language tasks. In parallel, recent studies in human neuroscience have begun positioning LLMs as computational models of human brain activity during context-rich, real-world language processing. In these works, researchers use encoding models to estimate a linear mapping between internal representations—namely, embeddings—extracted from an LLM and measurements of human brain activity, word by word during natural language comprehension. This simple approach of linearly ‘aligning’ the LLM’s internal feature space to human brain features has yielded remarkably good prediction performance in both functional magnetic resonance imaging (fMRI) and electrocorticography (ECoG). The high spatiotemporal resolution of invasive ECoG recordings, in particular, promises to provide finer-grained insights into shared representations and processes between LLMs and the brain. When exposed to the same natural language stimulus, such as a spoken story, human neural activity converges on stimulus features ranging from basic acoustic attributes to more complex linguistic and narrative elements. However, while a coarse alignment exists across individual brains, the finer cortical topographies for language representation exhibit notable idiosyncrasies among individuals. To address this, hyperalignment techniques have been developed in fMRI research to aggregate information across participants into a unified information space while overcoming the misalignment of functional topographies across participants. ECoG presents a more difficult correspondence problem than fMRI because each participant has a different number of electrodes in different locations (with placement guided by clinical considerations, not research goals). Thus, how to best aggregate electrodes across individuals is a matter of ongoing research. For this reason, encoding models are typically constructed separately at each electrode within individual participants and are not assessed for their generalization to new participants. In this Article, we measured the neural responses of eight ECoG participants implanted with invasive intracranial electrodes while they listened to a natural language stimulus. We develop a shared response model (SRM) to aggregate neural activity and isolate a stimulus-driven feature space that is shared across individuals. In parallel, we use LLMs to extract contextual embeddings for each word of the podcast. We then build encoding models to estimate a linear mapping from the contextual embeddings to the shared neural features. We show that the SRM yields substantially higher encoding performance than the original individual-specific electrodes. Moreover, we show that we can use this shared space to ‘denoise’ individual participant responses by projecting from the shared space back into the individual electrode space. We find that the SRM-reconstructed data yield the largest improvement in brain areas specialized for language comprehension. Finally, we demonstrate that the SRM allows us to construct encoding models that better generalize across participants.