Authors
Hilger, A. M., Soeding, J.
Abstract
Motivation: Multivariate survival analysis with hundreds of correlated outcomes is computationally challenging. Established approaches either ignore correlations between response variables, rely on black-box deep learning or are limited to small-scale outcomes. Results: We introduce MMAD-Risk, a novel multivariate mixed accelerated failure time (AFT) model that enables scalable analysis of high-dimensional survival analysis data. We train MMAD-Risk using amortized variational inference where we design the variational distribution such that it factorizes across diseases, allowing us to decompose multivariate disease risk prediction into a series of tractable, one-dimensional problems. This allows us to calculate the ELBO analytically, enabling fast computation. The model employs a low-rank decomposition of the effect size matrix B = VW to capture shared disease mechanisms and latent random effects Vz to model comorbidity. MMAD-Risk is trained on the UK Biobank Pharma Proteomics cohort (N {approx} 55000, P {approx} 3000 proteins, D = 271 diseases). Using the full 3,000-protein dataset, MMAD-Risk achieved a mean concordance index (c-index) of 0.744 for diagnoses occurring [>=] 10 years after blood sample collection, outperforming a Cox proportional hazards model (mean c-index = 0.709). Greedy backward selection identified a 10-protein panel that preserved > 99% of the full-model performance. On this reduced panel MMAD-Risk still outperformed Cox regression (0.738 vs. 0.679).
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 18 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 8
- Comments 0