Immune cell subsets linking insomnia trajectories and incident cognitive impairment: a prospective cohort study.
Authors: Chen L, Fu J, Jia X, Zhang L, Jin X, Wu S, Zhu B, Sun K, Fu D, Wang Y, Liu Z, Li S, Zhang J
Journal: BMC medicine
mental health
psychology
open access
Abstract
Cluster analysis is the task of partitioning a series of observations into several clusters (groups) in a way that the observations in the same cluster tend to be similar. Cluster analysis has been applied in various fields, from medicine to social sciences. Numerous methods have been developed to perform cluster analysis, such as k-means and hierarchical clustering [, ]. However, most basic clustering methods cannot be directly applied to real-world datasets because they often include missing values. Complete case analysis is the approach in which any observations with missing values are merely deleted from the analysis. Although this approach is one of the simplest ways to deal with missing values, it can influence the validity of results and impede subsequent analyses based on the assigned clusters to each observation. For example, if only a limited number of observations remain after the deletion of observations with missing values, the statistical power declines and the association between the assigned clusters and follow-up data might not be detectable. Therefore, we focus on the approaches that include all of the observations, even those with missing values. Multiple imputation is one of such methods [–]. This approach consists of two main steps: first, multiple complete datasets by imputing the missing values using probabilistic models and second, the results from each complete dataset into the final result using Rubin’s rule. While multiple imputation has been used for regression analysis with incomplete data, the approach has recently been extended to cluster analysis. For cluster analysis, instead of Rubin’s rule, cluster ensemble algorithms have been proposed to combine the results [, ]. Basagaña et al. [] employed a relabeling and voting algorithm to combine results from multiple complete datasets. Faucheux et al. [] used the MultiCons algorithm based on frequent item set mining []. Bruckers et al. and Audigier and Niang [, ] focused on the formulation of a clustering ensemble algorithm as the solution to the mean partition problem. They applied a direct optimization approach of the target function and the non-negative matrix factorization ensemble algorithm []. Other proposed methods to combine multiple imputation results are the incorporation of distance matrices and the integration of cluster centroids [, ].