Authors
Barbora Rehák Bučková, Charlotte Fraza, Cecilie Koldbæk Lemvigh, Camilla Bärthel Flaaten, Linn Sofie Sæther, Giulia Cattarinussi, Karen Marie Sandø Ambrosen, Ole Andreas Andreassen, Lars Tjelta Westlye, Christian Beckmann, Bjørn Hylsebeck Ebdrup, Paola Dazzan, Torill Ueland, Robin Mitra, Andre Marquand
Published in
Communications AI & computing. Volume 1. Issue 1. Pages 19. Epub Oct 09, 2026.
Abstract
Missing data are ubiquitous in biomedical research, and become particularly problematic when combining data from multiple clinical studies, an increasingly common practice for building models that generalise across populations. This process introduces complex structured patterns of missing data ('structured missingness'), where one study may omit variables that another collects, leaving deterministic gaps that standard imputation methods were not designed to handle. Here, we show that many popular imputation algorithms systematically fail under structured missingness, distorting the underlying statistical properties of the data in ways that conventional accuracy metrics cannot detect. Methods designed to preserve data distributions are more robust, and a hierarchical modelling approach that explicitly accounts for site-specific differences further improves performance. We introduce a multi-metric evaluation framework and a practical decision guide to support method selection. Our findings highlight the need for more principled imputation approaches in multi-site biomedical research and provide concrete tools to address this challenge.
PMID:
42859516
Bibliographic data and abstract were imported from PubMed on 11 Oct 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 1
- Comments 0