← Back to Research Papers

"What bothers the heart": a phenomenological study of adolescents' experiences and interpretations of school-based gender-based violence in Kinshasa, DRC.

Authors: Hernandez JH, Mbadu MF, Murrairi L
Journal: Sexual and reproductive health matters
mental health psychology open access

Abstract

Characterizing genetic relationships within single-cell transcriptomic data [,], and leveraging these representations for a broad spectrum of downstream biological tasks, has emerged as a central paradigm across modern computational biology and related disciplines [–]. To this end, a wide range of computational approaches has been developed to systematically learn latent genetic structures from such data and generalize these representations to predictive tasks, including perturbation response modeling, drug sensitivity prediction, disease classification, and cellular dynamics inference [–]. Early efforts adapted classical frameworks, such as multilayer perceptrons (MLPs) and differential expression testing (DET) [], while subsequent work introduced specialized models tailored to single-cell data, including scVI, scANVI, and GEARS [–], collectively demonstrating the feasibility of extracting biological representations from sparse transcriptomic profiles. More recently, single-cell transcriptomic foundation models have been proposed as a unified representational framework, leveraging self-attention mechanisms [] to capture complex gene–gene dependencies and enabling flexible adaptation to a wide range of downstream tasks through pretraining and fine-tuning. Emerging models such as scGPT, scFoundation, and SCmilarity [,,] highlight the potential of this paradigm. In parallel, the Bidirectional Encoder Representations from Transformers (BERT) architecture [] has been adapted to single-cell transcriptomics. An early example, scBERT [], applied a BERT-style framework to single-cell data and demonstrated its utility for cell-type annotation, while the subsequent Geneformer (GF), a ranking-based BERT variant pretrained on approximately 30 million cells [], extended this paradigm toward foundation-model applications across a broader range of downstream tasks. Despite these advances, foundation models such as GF remain susceptible to fundamental mismatches between backbone design choices and the intrinsic ranking properties of single-cell transcriptomic data, an issue that has yet to be systematically examined. In particular, GF operates on rank-ordered gene expression profiles, yet directly inherits the masked language modeling (MLM) objective from natural language processing [], where token repetition is permissible. This misalignment introduces undesirable behaviors: GF can repeatedly predict the same gene within a single cellular profile, despite the biological constraint that each gene is uniquely represented per cell. Consequently, the model exhibits a bias toward highly frequent and ubiquitously expressed genes, leading to reduced prediction diversity and diminished biological interpretability. In addition, recent studies in foundation models have highlighted that indiscriminate scaling of pretraining data does not necessarily yield improved generalization performance [–]. Within the context of single-cell transcriptomic modeling, this issue remains largely unexplored. Current approaches predominantly emphasize increasing dataset size, often aggregating ever-larger collections of cells, yet lack a principled assessment of whether such scaling is necessary or effective for representation learning in this domain. To address these underexplored limitations, we introduce GF, a modified GF architecture designed to better align model behavior with the rank-ordered structure of single-cell transcriptomic data. Central to this design is a cumulative assignment and balancing (CAB) module, implemented as a post-prediction processor. The CAB module incorporates a probability cumulative-assignment mechanism to propagate positional constraints across predictions, ensuring consistency with the non-redundant nature of gene rankings within each cell. In parallel, a similarity-based regularization term penalizes redundant or overly similar probability distributions across gene positions, thereby promoting diversity in predicted gene identities. Together, these components explicitly encode structural priors of the data into the prediction process, mitigating the mismatch introduced by conventional masked language modeling objectives. To further investigate the role of data scale in conjunction with architectural design, we additionally construct a reduced pretraining corpus (Genecorpus-1M) by uniformly subsampling one million profiles from the original Genecorpus-30M, enabling controlled comparisons across different data regimes ().