Authors
Gaetan Kamdje Wabo, Piotr Pawel Sokolowski, Mahboubeh Jannesari Ladani, Michael Hagmann, Thomas Ganslandt, Fabian Siegel
Published in
JMIR medical informatics. Jul 08, 2026. Epub Jul 08, 2026.
Abstract
Secondary use of electronic health record data requires robust privacy protection. k-Anonymity is widely used to enable data sharing by ensuring that each quasi-identifier combination occurs in at least k records, yet its analytical impact on clinically meaningful structures remains insufficiently characterized, particularly for the combination of record suppression and microaggregation that arises when a numeric attribute lacks a natural generalization hierarchy. A further gap is that anonymization tools report internal information-loss values but do not signal the downstream distributional and inferential distortions these transformations introduce.
This study evaluated the analytical footprint of k-anonymity at k = 5, 10, and 15 on two core data elements in retrospective hospital research: primary ICD-10-GM diagnosis codes and hospital length of stay (LOS). It aimed to determine and quantify whether anonymization introduces meaningful distortions not captured by the anonymization tool itself, and whether diagnosis-specific LOS patterns remain reproducible after anonymization.
We analyzed 719,387 inpatient encounters from University Hospital Mannheim from 2010 to 2024. Anonymization was performed with the ARX tool. It used record suppression and microaggregation. Distributional distortion was assessed with the Kolmogorov-Smirnov (KS) D statistic, quantile shifts, interquartile range changes, and tail changes. Categorical fidelity was quantified using the Jaccard coefficient and Cramer's V. Inferential reproducibility was evaluated based on a three-level linear mixed model. The model included random intercepts for ICD-3 and patient. In addition, we investigated the intraclass correlation coefficient (ICC) and diagnosis-level effect concordance. Moreover, the concordance was measured using Spearman rho and Lin's concordance correlation coefficient (CCC), both with 95% confidence intervals. A composite traffic-light verdict summarized the results.
ARX masked quasi-identifier cells rather than deleting rows; the proportion of encounters with a masked cell rose from 0.77% (k=5) to 2.62% (k=15), distributed almost uniformly across admission years. KS D was stable at 0.147. Median LOS shifted one day, standard deviation declined by 6.5 days, and diagnosis-level mean fell 1.81 days (~28% of 6.52-day baseline). Jaccard overlap fell from 0.624 to 0.421. The diagnosis ICC rose from 0.294 to 0.837, reflecting variance compression rather than improved signal. Best linear unbiased prediction (BLUP) rank concordance (Spearman rho 0.964-0.970) and aggregate magnitude agreement (Lin CCC 0.959-0.966) were high, yet about 7% of low-signal diagnoses showed sign reversals. Most distributional change occurred at k=5.
k-Anonymity preserved the ranking of diagnosis-specific LOS effects but altered distributional shape, individual effect magnitudes, and diagnostic vocabulary, none of which was flagged by the internal loss metric. Anonymized data of this type may support ordinal analyses but can mislead analyses requiring faithful variance structure, accurate absolute effects, or complete rare-diagnosis representation. A reporting checklist is provided to document these effects.
PMID:
42593356
Bibliographic data and abstract were imported from PubMed on 13 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 1
- Comments 0