Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Detecting and Mitigating AI Bias in Health Care: Development and Validation of a Unified Multistage Framework.

Created on 15 Aug 2026

Authors

Ruj Mateedulsatit, Chetneti Srisa-An

Published in

JMIR AI. Volume 5. Pages e102146. Aug 14, 2026. Epub Aug 14, 2026.

Abstract

AI-driven clinical systems can improve diagnosis, prognosis, and resource allocation, but they may reproduce disparities encoded in historical health care data. Existing mitigation methods typically target a single source of bias, while clinical datasets often contain interacting representation, proxy, integrity, and temporal biases.
This study aims to develop and systematically evaluate a prespecified multistage workflow for detecting representation, missingness, proxy, integrity, and temporal biases and model performance limitations in structured health care datasets; apply prespecified mitigation actions when audit criteria are met; and determine whether these actions improve predictive discrimination and demographic fairness compared with a conventional random forest baseline.
We performed a fresh, deterministic reconstruction from the raw public data, using a patient-level 80/20 split for Diabetes 130-US Hospitals. M2 was a conventional random forest with median or mode imputation and training-only categorical encoding. M3 added explicit missingness features and poststratification weights clipped at the 95th percentile. Race was excluded from prediction and used for auditing and weighting. CMS SynPUF was modeled separately for a compatible claims-based readmission task; the National Health and Nutrition Examination Survey was limited to stage-level representation, proxy, missingness, and bounded-laboratory audits. Five hundred stratified bootstrap replicates were used for overall metrics, and 300 were used for subgroup metrics.
The Diabetes test set contained 20,203 encounters from 14,304 patients, including 2254 (11.16%) positive outcomes. M3 improved the macro F1 from 0.514 to 0.546, reduced the Brier score from 0.231 to 0.213, and reduced the demographic parity difference from 0.206 to 0.124, but the area under the receiver operating characteristic curve (AUC) decreased from 0.648 to 0.640, and the equalized odds ratio decreased from 0.444 to 0.291. In CMS SynPUF (12,801 test episodes; 1232 positives), the AUC was similar (0.798 vs 0.796) and the Brier score improved slightly (0.187 vs 0.183), whereas the demographic parity difference increased from 0.693 to 0.730. The exploratory rule-gated mixture of experts did not improve fairness, the Brier score, or the macro F1 relative to M3. Detailed subgroup performance, calibration, missingness analyses, and bootstrap CIs were comprehensively evaluated in this study.
In this reproducible retrospective reconstruction, missingness-aware weighting improved the macro F1, Brier score, and demographic parity on the primary test set but did not improve the AUC or equalized odds ratio. Centers for Medicare and Medicaid Services results did not reproduce a fairness improvement, and the exploratory mixture of experts did not outperform the M3 quality expert on most outcomes. The findings demonstrate a fairness-calibration-discrimination trade-off rather than uniform improvement and do not establish clinical deployment readiness.

PMID:
42600144
Bibliographic data and abstract were imported from PubMed on 15 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 3
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement