← Back to Research Papers

Quantifying uncertainty of predictions from cancer progression models.

Authors: Hu YL, Pfahler S, Lösch A, Vocht S, Hansch S, Rupp K, Beerenwinkel N, Wettig T, Schill R, Spang R
Journal: Bioinformatics (Oxford, England)
mental health psychology open access

Abstract

Cells are the fundamental units of biology, and many other biological entities are interpreted in relation to specific cell types [,]. The importance of cell types is particularly evident in biological processes and disease states governed by interactions among diverse cellular and molecular components [–]. The computational task of identifying cell types in published literature is important for advancing biomedical research and healthcare-related computational modeling. The need for accurate cell-type extraction is illustrated by the tumor microenvironment, where interactions among diverse cell types, including tumor cells, immune cells, and stromal cells, influence disease progression and treatment response [–]. These interactions can be modeled using quantitative systems pharmacology (QSP) approaches to predict tumor progression and treatment response [–]. The predictive value of such models depends on the accuracy and completeness of the underlying tumor microenvironment interaction network. Because these networks are often constructed from biomedical literature, reliable cell-type extraction is an important first step toward building more accurate interaction networks and supporting downstream computational modeling [,,]. The task of extracting cell-type entities from literature has two linked parts: (i) named entity recognition (NER), which identifies entities in text, such as cell types or cell populations, and (ii) named entity normalization (NEN), also called entity linking, which maps each entity to a standard identifier. The entity-normalization step is needed because the same concept can appear in many forms in text. For example, a cell type may be represented by alternative names or abbreviations, such as “granulosa cell” or “GC,” and “endothelial cell” or “endotheliocyte.” Normalization maps these forms to the same Cell Ontology (CL) identifier [], which improves consistency across documents and supports downstream analysis. Current off-the-shelf biomedical text-mining tools offer useful capabilities, but their support for cell-type recognition and normalization remains uneven. PubTator3 [] recognizes genes/proteins, chemicals, diseases, species, variants, and cell lines, but not cell types. HunFlair2 [] supports cell lines, chemicals, diseases, genes, and species but does not target cell types. GNorm2 [], a strong system for gene recognition and normalization, is limited to genes. VANER2 [] expands biomedical NER coverage, but it is a recognition system and does not link entities to ontology identifiers. Although the more comprehensive BERN2 system [] includes both cell-type recognition and normalization, its normalization strategy relies on rule-based methods, which have limited coverage. The scispaCy toolkit [] provides biomedical NER models with cell-related labels. However, its built-in linker does not support the Cell Ontology. As a result, an external integration such as PyOBO [] is needed for CL linking. Here, we present an informatics tool that addresses an unmet need for practical, end-to-end support for both cell-type recognition and Cell Ontology normalization.