← Back to Research Papers

Peripheral blood DNA methylation partially reports on immune cell type composition in the frontal brain.

Authors: Meijer M, Fu MPY, Navarro-Delgado EI, Engelbrecht HR, Turecki G, Chan MH, Kobor MS
Journal: Brain, behavior, & immunity - health
mental health psychology open access

Abstract

In many biomedical applications, the main interest centers around contrastive analysis, targeting variation in the sample with respect to a reference group. Contrasted groups may be disease versus disease-free, treatment versus control, measurements within the same sample before versus after treatment. In such contrastive scenarios with data from more than one group, the primary goal is to separate and identify variation patterns shared across groups versus those that are unique to a particular group. Despite their natural appeal in biomedical applications, few methods have been developed to specifically explore and take advantage of these contrastive setups. To address this limitation, Abid et al. [] proposed the contrastive Principal Component Analysis (cPCA) for multivariate data. cPCA aims to highlight unique variation patterns in a target data set that are different from those shared with a reference group. One example would be targeting unique variations in a treatment group beyond those which may be viewed as common sample variation explained by age, sex, and other demographic variables. Hence, by removing shared variation patterns with the control group, cPCA highlights variation within the treatment group that are beyond those that can be viewed as natural sample variation due to demographic characteristics as observed within the control sample. cPCA computes multiple subspaces by subtracting varying amounts of background variation from the target variation and calculates the dominant principal components (PCs) within each subspace, referred to as contrastive PCs. Building on this concept, Severson et al. [] proposed contrastive Latent Variable Models (cLVM), which further delineate the underlying contrastive structure using latent variables that are unique or shared across two samples. cLVM offers a more straightforward and comprehensive approach to contrastive modeling, without reliance on additional tuning parameters such as the amount of background variation subtracted from the target variation. Contrastive inference is also of great interest in functional data analysis, where functional data is collected in more than one group and there is interest in identifying time-dynamic shared and unique variation patterns across groups. The literature on functional data analysis has grown rapidly over the past two decades, providing efficient and highly-interpretable representations of data in the form of curves, trajectories, or surfaces observed over continuous domains [, , , ]. Functional Principal Component Analysis (FPCA) emerged as a fundamental modeling tool in this literature, providing dimension reduction and capturing dominant modes of variation from complex data sets [, ]. FPCA has been adapted for modeling of functional observations with temporal and spatial correlation structures [, , , ], as well as sparse functional observations, multilevel structures, high-dimensional functional outcomes and multivariate functional data [-, , , , , , , , ]. Latent factor models have also been utilized in reconstruction of functional trajectories and exploration of the underlying correlation structures in lower dimensions [, ]. Li and Xiao [] extended latent factor models to multivariate functional data with the goal of dimension reduction and ease in interpretation of the parsimonious shared latent processes underlying multivariate functional data. With the goal of characterizing unique variation in the target group, Zhang and Li [] extended cPCA to functional data settings (CFPCA) via relying on the difference of target and background functional variation. However, developments depend on the choice of the amount of background variation subtracted from target variation, similar to cPCA and do not allow for unique variation in both data sets (target and background). . Different from recent formulations of CFPCA [], cLFM does not rely on the tuning parameter of the amount of background variation subtracted from target variation and allows for unique variation in both data sets. The first two applications are on contrastive scenarios involving electroencephalography (EEG) data collected on children diagnosed with autism and their neurotypical peers, while the third example examines kidney decline among mild to severe albuminuria subgroups within patients with chronic kidney disease (CKD).