Authors
F. Stacchietti, M. Nicolini, L. Chimirri, P.N. Robinson, E.Casiraghi, G. Valentini
Published in
Neurocomputing. Volume 702. Issue 134682. 2026.
Abstract
Mendelian genetic diseases comprise approximately 10,000 described disorders, yet the genetic basis is known only for about half of them, and a molecular diagnosis often remains difficult or unresolved. In this context, machine learning methods play a key role. However, identifying pathogenic variants in non-coding regions of the human genome is particularly challenging, as they are vastly outnumbered by neutral variants. This extreme imbalance causes standard machine learning approaches to exhibit a strong predictive bias towards the majority class, significantly limiting their sensitivity.
Building on recent advances in imbalance-aware and ensemble learning methods, we propose two novel deep learning models for predicting pathogenic non-coding variants in Mendelian diseases.
The first model, T-ResNet (Tabular Residual Neural Network), adopts a modular architecture with residual connections, along with a mini-batch balancing strategy to address class imbalance. This design simplifies hyperparameter optimization while mitigating vanishing-gradient effects. The second model, TIDE-Var (Tabular Implicit Deep neural network Ensembles for Variant prediction), leverages the TabM and BatchEnsemble models to build an implicit ensemble of deep neural networks trained jointly by minimizing a common objective function, and partially sharing learning parameters. Learner-specific adapter parameters promote diversity among the implicit base learners while keeping the ensemble computationally efficient.
Genome-wide experiments show that T-ResNet achieves an average Area Under the Precision-Recall Curve (AUPRC) of 0.630 ±0.033, comparable to that of HyperSMURF, a state-of-the-art method. In contrast, TIDE-Var yields significantly better results than HyperSMURF (AUPRC =0.714 ±0.025) and slightly better than XGBoost, one of the top-methods for the classification of tabular data. Ablation studies confirm that regularized joint ensemble learning with partially shared and base-learner specific adapter learning parameters are key factors in achieving high predictive performance, essential to improve the diagnostic yield for patients with rare genetic diseases.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 11
- Comments 0