Authors
Roxane Girault, Jassim Bensafir, Antoine Lamer, Jean-Baptiste Beuscart, Michaël Génin, Amadou Tidiane Niang, Slim Hammadi, Emmanuel Chazard
Published in
International journal of medical informatics. Volume 222. Pages 106711. Sep 03, 2026. Epub Sep 03, 2026.
Abstract
Although data reuse is increasingly common in healthcare research, changing regulatory frameworks are impeding these efforts. Synthetic data constitute a promising issue but also create challenges, such as the assessment of data fidelity and the ability to output identical statistical results. Conventional validation approaches often rely on univariate or bivariate comparisons, which fail to capture the complexity of associations between multivalued categorical variables.
We built and studied a fictitious database of 10,000 hospital stays reproducing the structure of the French Programme de Médicalisation des Systèmes d'Information database. Each stay included single-valued variables (one value per individual: sex, age in deciles, and diagnosis-related group) and multivalued variables (zero, one or several values per individual: diagnoses coded according to the International Classification of Diseases, 10th Edition, and procedures coded according to the French Classification Commune des Actes Médicaux).
All categorical variables were binarized, and thousands of pairwise association metrics (primarily odds ratios) were calculated for the reference and evaluation datasets. The results were summarized using curves, bubble charts, heatmaps, and coefficients such as exponential mean deviation. Simulated data degradations from 0% to 100% were introduced to evaluate the method's sensitivity.
We analyzed 500 ICD-10 diagnoses and 450 CCAM procedures, representing 225,000 possible combinations. We developed and evaluated graphical representations for assessing data fidelity at a glance. In simulations of an increasing degree of data degradation, those graphical representations and comprehensive, quantitative metrics facilitated the detection of the gradual loss of data fidelity.
We developed a simple, scalable, agnostic framework for assessing the fidelity of healthcare databases by systematically analyzing associations among the modalities of coded variables. This method complements existing approaches. It is particularly suitable for the evaluation of synthetic relational databases because it offers both general and granular insights into data fidelity loss.
PMID:
42700765
Bibliographic data and abstract were imported from PubMed on 06 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 2
- Comments 0