Authors
Chenggong Xie, Weichang Kong, Liuting Pi, Diaoxin Qi, Yuxuan Yang, Bin Wang, Hao Liang
Published in
Journal of evidence-based medicine. Pages e70166. Jul 25, 2026. Epub Jul 25, 2026.
Abstract
To systematically evaluate the diagnostic performance of large language models (LLMs) in automated medical literature screening and to determine their potential role in supporting evidence synthesis workflows.
PubMed, Web of Science, Embase, The Cochrane Library, Google Scholar, CNKI, Wanfang, VIP, and CBM were searched from January 1, 2022 to June 11, 2026. Studies assessing LLMs for automated title and abstract screening or full-text eligibility assessment in medical literature were included. The primary outcomes were sensitivity and specificity. Secondary outcomes included positive and negative likelihood ratios, diagnostic odds ratio, area under the curve (AUC), and efficiency-related metrics. Pooled sensitivity and specificity were estimated using a bivariate random-effects model and hierarchical summary receiver operating characteristic framework. Subgroup analyses and meta-regression were performed to explore sources of heterogeneity. This systematic review was registered with the Open Science Framework (https://osf.io/56d8j).
Eighteen studies published between 2023 and 2025 were included. In title and abstract screening, the pooled sensitivity was 0.92 (95% confidence interval [CI]: 0.81-0.96) and pooled specificity was 0.94 (95% CI: 0.90-0.97). The summary receiver operating characteristic AUC reached 0.98 (95% CI: 0.96-0.99). In full-text screening, pooled sensitivity and specificity both reached 0.99 (95% CI: 0.95-1.00) and the AUC was 0.99 (95% CI: 0.98-1.00). Prompt strategies incorporating examples or chain-of-thought reasoning were associated with higher sensitivity than strategies without these approaches (0.95 vs. 0.86, p < 0.01). Several studies reported substantial efficiency gains, including workload reductions ranging from approximately 50% to 99% and screening time reductions of up to tenfold.
LLMs shows promising performance in automated medical literature screening, particularly in full-text assessment. These models show strong potential as high sensitivity assistive tools that can substantially reduce manual screening burden while supporting evidence synthesis. Further real-world validation is needed to establish their role in evidence-based medicine.
PMID:
42499245
Bibliographic data and abstract were imported from PubMed on 25 Jul 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 4
- Comments 0