Authors
Cai, H., Xu, K., Osmundson, S., Ajilore, O., Andreescu, C., Taylor, W. D., Liu, J., Zhang, X.
Abstract
The rapid growth of technologies such as wearable devices and high-throughput sequencing has generated unprecedented volumes of longitudinal data, often with multiple correlated features. This trend highlights the need for robust clustering methods tailored to multi-feature longitudinal datasets. Although previous reviews have benchmarked clustering methods in this context, most evaluations relied heavily on simulated and overly simplified datasets, which substantially limit their interpretability and applicability to real-world data. Missing data is also ubiquitous in longitudinal data, yet understudied in the context of clustering. To address these gaps, we systematically evaluate clustering methods across three distinct types of real-world longitudinal datasets: (1) a high-dimensional, low-timepoint metagenomic study, (2) a low-dimensional, high-timepoint multi-site repeated-measure study, and (3) a high-dimensional, high-timepoint repeated-measure study. We further investigate how missing data and different imputation strategies influence clustering performance. In addition, we introduce a novel resampling framework based on the empirical cumulative distribution function (eCDF) and a copula. This framework enables us to simulate multi-feature longitudinal data that preserves the correlation structure of real datasets, while also allowing for controlled variations such as sample size. Our simulation framework naturally accommodates missing data, thereby providing a flexible tool for benchmarking and comparing clustering methods in more realistic longitudinal settings.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 02 Oct 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 10
- Comments 0