Authors
Ziming Luo, Atoosa Kasirzadeh, Nihar B Shah
Published in
Proceedings of the National Academy of Sciences of the United States of America. Volume 123. Issue 41. Pages e2610214123. Oct 13, 2026. Epub Oct 05, 2026.
Abstract
The rapid rise of autonomous AI scientists marks a paradigm shift in scientific discovery by automating the research lifecycle. Yet their rushed development has outpaced critical oversight, leaving key workflow decisions dangerously unscrutinized. We present a much-needed systematic analysis of open-source AI scientist systems, investigating four primary pitfalls: inappropriate benchmark selection, data leakage, metric misuse, and post hoc selection bias. Through controlled experiments that isolate each pitfall, we find systematic vulnerabilities across two representative open-source systems. Crucially, we find that these flaws are largely invisible at the level of the final manuscript, suggesting that current manuscript-centric peer review paradigms are fundamentally insufficient for ensuring the integrity of automated research. We further propose mitigation strategies and demonstrate that access to full workflow artifacts (log traces and code) enables more effective auditing. Our findings suggest that journals, conferences, and researchers should move beyond manuscript-only evaluation toward process auditing the end-to-end workflow artifacts of AI scientist systems.
PMID:
42832657
Bibliographic data and abstract were imported from PubMed on 06 Oct 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 52
- Comments 0