Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Artificial Intelligence-Assisted Data Extraction With a Large Language Model: A Study Within Reviews.

Created on 04 Nov 2025

Authors

Gerald Gartlehner, Shannon Kugley, Karen Crotty, Meera Viswanathan, Andreea Dobrescu, Barbara Nussbaumer-Streit, Graham Booth, Jonathan R Treadwell, Jung Min Han, Jesse Wagner, Eric A Apaydin, Erin L Coppola, Margaret Maglione, Rainer Hilscher, Robert Chew, Meagan Pilar, Bryan Swanton, Leila C Kahwati

Published in

Annals of internal medicine. Nov 04, 2025. Epub Nov 04, 2025.

Abstract

Data extraction is a critical but error-prone and labor-intensive task in evidence synthesis. Unlike other artificial intelligence (AI) technologies, large language models (LLMs) do not require labeled training data for data extraction.
To compare an AI-assisted versus a traditional, human-only data extraction process.
Study within reviews (SWAR) using a prospective, parallel-group comparison with blinded data adjudicators.
Workflow validation within 6 ongoing systematic reviews of interventions under real-world conditions.
Initial data extraction using an LLM (Claude, versions 2.1, 3.0 Opus, and 3.5 Sonnet) verified by a human reviewer.
Concordance, time on task, accuracy, sensitivity, positive predictive value, and error analysis.
The 6 systematic reviews in the SWAR yielded 9341 data elements from 63 studies. Concordance between the 2 methods was 77.2% (95% CI, 76.3% to 78.0%). Compared with the reference standard, the AI-assisted approach had an accuracy of 91.0% (CI, 90.4% to 91.6%) and the human-only approach an accuracy of 89.0% (CI, 88.3% to 89.6%). Sensitivities were 89.4% (CI, 88.6% to 90.1%) and 86.5% (CI, 85.7% to 87.3%), respectively, with positive predictive values of 99.2% (CI, 99.0% to 99.4%) and 98.9% (CI, 98.6% to 99.1%). Incorrect data were extracted in 9.0% (CI, 8.4% to 9.6%) of AI-assisted cases and 11.0% (CI, 10.4% to 11.7%) of human-only cases, with corresponding proportions of major errors of 2.5% (CI, 2.2% to 2.8%) versus 2.7% (CI, 2.4% to 3.1%). Missed data items were the most frequent error type in both approaches. The AI-assisted method reduced data extraction time by a median of 41 minutes per study.
Assessing concordance and classifying errors required subjective judgment. Consistently tracking time on task was challenging.
Data extraction assisted by AI may offer a viable, more efficient alternative to human-only methods.
Agency for Healthcare Research and Quality and RTI International.

PMID:
41183336
Bibliographic data and abstract were imported from PubMed on 04 Nov 2025.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 83
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement