Authors
Amalie Koch Andersen, Hadi Mehdizavareh, Arijit Khan, Tobias Becher, Simone Britsch, Markward Britsch, Morten Bøttcher, Simon Winther, Palle Duun Rohde, Morten Hasselstrøm Jensen, Simon Lebech Cichosz
Published in
Journal of the American Medical Informatics Association : JAMIA. Jul 30, 2026. Epub Jul 30, 2026.
Abstract
Machine-learning-based clinical risk prediction models are increasingly used to support decision-making in healthcare. While class-imbalance correction techniques are commonly applied to address rare outcomes, their impact on probabilistic calibration remains insufficiently understood. This study evaluated the effect of widely used resampling strategies on both discrimination and calibration across real-world clinical prediction tasks.
Ten clinical datasets spanning diverse medical domains and including over 600 000 patients were analyzed. Multiple machine-learning model families were evaluated. Models were trained on original data and using 3 1:1 class-imbalance correction strategies (synthetic minority oversampling technique, random undersampling, and random oversampling). Performance was assessed on held-out data using discrimination and calibration metrics.
Resampling had no positive impact on predictive performance. Changes in area under the receiver operating characteristic curve (ROC-AUC) and precision-recall AUC were small and inconsistent (ROC-AUC: -0.002 to -0.01; PR-AUC: -0.10 to -0.03), with no method showing systematic improvement. In contrast, calibration was consistently degraded. Resampled models showed higher Brier scores (increase 0.029-0.080) and marked deviations in calibration intercept and slope, indicating distorted predicted risks despite preserved ranking performance.
Across diverse clinical datasets, resampling primarily altered the implicit class prior learned during training, leading to miscalibration when models were evaluated. The consistent dissociation between discrimination and calibration highlights that rank-based metrics alone are insufficient for evaluating clinical utility. Gains from imbalance correction can typically be reproduced by threshold adjustment without distorting predicted probabilities.
Common 1:1 class-imbalance correction techniques do not improve discrimination and may substantially degrade calibration, limiting their suitability for clinical risk prediction where accurate probabilities are essential.
PMID:
42533626
Bibliographic data and abstract were imported from PubMed on 31 Jul 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 10
- Comments 0