Authors
Kuete Fouodo, C. J., Kapar, J., Huels, A., Liang, D., Wright, M. N.
Abstract
Data availability is critical for understanding complex disease pathways and developing robust predictive models. Although high-throughput omics technologies have improved insight into disease mechanisms, data acquisition from inaccessible tissues such as the central nervous system remains a major limitation, causing small sample sizes and complicating early prediction of neurodegenerative disorders such as Alzheimer's and Parkinson's diseases. Generative modeling has emerged as a powerful approach for synthesizing data to support downstream clustering and prediction with small sample size, but existing methods rarely handle high-dimensional tabular omics data effectively. Adversarial random forests (ARFs) provide a well-performing framework for tabular data generation but are not designed for high-dimensional settings. To address this limitation, we introduce high-dimensional ARF (h-ARF), an extension of ARF optimized for integrated clinical and high-dimensional omics data. Using benchmarks across nine datasets and eight performance metrics, we show that h-ARF better preserves both feature distributions, and downstream clustering and prediction utilities compared with ARFs. The method is implemented in the open-source R package harf, available on CRAN.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 16 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 9
- Comments 0