Tipping the balance: impact of class imbalance correction on the performance of clinical risk prediction models : research and applications
Machine-learning-based clinical risk prediction models are increasingly used to support decision-making in healthcare. While class-imbalance correction techniques are commonly applied to address rare outcomes, their impact on probabilistic calibration remains insufficiently understood. This study ev...
Salvato in:
| Autori principali: | , , , , , , , , , , |
|---|---|
| Natura: | Article (Journal) |
| Lingua: | inglese |
| Pubblicazione: |
30 July 2026
|
| In: |
Journal of the American Medical Informatics Association
Year: 2026, Pages: 1-14 |
| ISSN: | 1527-974X |
| DOI: | 10.1093/jamia/ocag127 |
| Accesso online: | Verlag, lizenzpflichtig, Volltext: https://doi.org/10.1093/jamia/ocag127 |
| Note sull'autore: | Amalie Koch Andersen, MSc, Hadi Mehdizavareh, MSc, Arijit Khan, PhD, Tobias Becher, MD, Simone Britsch, MD, Markward Britsch, PhD, Morten Bøttcher, MD, PhD, Simon Winther, MD, PhD, Palle Duun Rohde, PhD, Morten Hasselstrøm Jensen, PhD, Simon Lebech Cichosz, PhD |
| Riassunto: | Machine-learning-based clinical risk prediction models are increasingly used to support decision-making in healthcare. While class-imbalance correction techniques are commonly applied to address rare outcomes, their impact on probabilistic calibration remains insufficiently understood. This study evaluated the effect of widely used resampling strategies on both discrimination and calibration across real-world clinical prediction tasks.Ten clinical datasets spanning diverse medical domains and including over 600 000 patients were analyzed. Multiple machine-learning model families were evaluated. Models were trained on original data and using 3 1:1 class-imbalance correction strategies (synthetic minority oversampling technique, random undersampling, and random oversampling). Performance was assessed on held-out data using discrimination and calibration metrics.Resampling had no positive impact on predictive performance. Changes in area under the receiver operating characteristic curve (ROC-AUC) and precision-recall AUC were small and inconsistent (ROC-AUC: −0.002 to −0.01; PR-AUC: −0.10 to −0.03), with no method showing systematic improvement. In contrast, calibration was consistently degraded. Resampled models showed higher Brier scores (increase 0.029-0.080) and marked deviations in calibration intercept and slope, indicating distorted predicted risks despite preserved ranking performance.Across diverse clinical datasets, resampling primarily altered the implicit class prior learned during training, leading to miscalibration when models were evaluated. The consistent dissociation between discrimination and calibration highlights that rank-based metrics alone are insufficient for evaluating clinical utility. Gains from imbalance correction can typically be reproduced by threshold adjustment without distorting predicted probabilities.Common 1:1 class-imbalance correction techniques do not improve discrimination and may substantially degrade calibration, limiting their suitability for clinical risk prediction where accurate probabilities are essential. |
|---|---|
| Descrizione del documento: | Gesehen am 18.08.2026 |
| Descrizione fisica: | Online Resource |
| ISSN: | 1527-974X |
| DOI: | 10.1093/jamia/ocag127 |