Beyond Area Under the Receiver Operating Characteristic Curve: Evaluating Predictive Performance Metrics Under Class Imbalance in Real-World Clinical Data.

Publication date: Jun 24, 2026

Predictive models increasingly support clinical decision-making, although imbalanced outcome distributions are common in health care datasets and can distort performance evaluation. The area under the receiver operating characteristic curve (AUROC) remains the most frequently reported metric, despite its limited ability to reflect clinically meaningful performance under class imbalance. This study aimed to examine the influences of metric selection on the clinical interpretation of predictive models in imbalanced real-world health care data. This was a retrospective cohort study, including 17,018 hospitalized patients with COVID-19. Two predictive models using extreme gradient boosting (XGBoost) were developed to predict kidney replacement therapy (KRT) and mortality. Model performance was assessed using AUROC, macro-F1-score, class-specific precision and recall, calibration (curve, slope, and intercept), decision curve analysis, and learning curves. Standard rebalancing strategies were applied exclusively to the training data to evaluate their impact on performance. KRT occurred in 9. 5%, and mortality in 18. 0%. Although AUROC values were high (0. 928 for KRT and 0. 945 for mortality), performance in the minority class was substantially lower. For KRT, precision was 0. 539 and recall 0. 372; for mortality, precision was 0. 725 and recall 0. 718. Rebalancing strategies were associated with higher recall for the minority class, but this gain was accompanied by a reduction in precision, with minimal impact on AUROC values. As a result, AUROC remained high despite clinically relevant changes in error distribution between false positives and false negatives. The learning curves show a plateau-like shape, with stable validation performance across all training set sizes for both outcomes. AUROC alone is insufficient to evaluate prediction models in imbalanced health care scenarios, even with rebalancing. Routine reporting of class-aware metrics, alongside learning curve analysis, is essential to support robust and clinically meaningful evaluation of predictive models, rather than their direct translation into practice.

Open Access PDF

Concepts Keywords
Covid Area Under Curve
F1 artificial intelligence
Kidney F-score
Mortality Female
Rebalancing Humans
learning curve
Male
performance metrics
Prediction Algorithms
Predictive Learning Models
predictive model
Retrospective Studies
ROC Curve

Semantics

Type Source Name
disease MESH COVID-19
drug DRUGBANK Flunarizine
pathway REACTOME Translation
drug DRUGBANK Etodolac
drug DRUGBANK Ribostamycin
disease MESH Neglected Diseases
drug DRUGBANK Coenzyme M
disease MESH death
disease MESH cardiovascular disease
drug DRUGBANK Trestolone
disease MESH included
disease MESH ers
drug DRUGBANK Methionine
disease MESH tics
drug DRUGBANK Isoxaflutole
drug DRUGBANK Aspartame
disease MESH confusion
disease MESH dis
disease MESH char
disease MESH strain
disease MESH emergency
disease MESH ces
disease MESH thrombosis
drug DRUGBANK Etoperidone
drug DRUGBANK Dimethyl sulfone
disease MESH BPP
disease MESH CAP
disease MESH osteoarthritis
disease MESH ACC
drug DRUGBANK Acetohydroxamic acid
disease MESH dementia
drug DRUGBANK Saquinavir
pathway REACTOME Reproduction

Original Article

(Visited 1 times, 1 visits today)

Leave a Comment

Your email address will not be published. Required fields are marked *