Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Benchmarking a Local Schema-Constrained Large Language Model Pipeline for Abstract Screening and Evidence Mapping.

Created on 22 Jul 2026

Authors

Alessandro Serretti

Published in

Cureus. Volume 18. Issue 6. Pages e111193. Epub Jun 20, 2026.

Abstract

Background The growth of biomedical literature increasingly exceeds the capacity of manual evidence synthesis. Large language models (LLMs) may support abstract screening and structured extraction, but many current workflows depend on proprietary cloud APIs, creating challenges for governance, reproducibility, and scalable deployment. Methods I developed a fully local, open-weight, schema-constrained pipeline (gpt-oss-20b, deployed via Ollama on Apple M1 Max) for title/abstract-based scoping workflows. The pipeline combined deterministic metadata filtering, LLM-assisted screening, and structured abstract extraction. Performance was benchmarked against three published systematic reviews (ketamine/neuroimaging; clozapine/suicidality; clozapine patient/caregiver perspectives) using precision, recall, and F1 against reference inclusion sets. I also report audit-adjusted estimates (i.e., performance metrics recalculated after manual full-text adjudication of discrepant records) alongside standard reference-set performance. Results In the ketamine/neuroimaging benchmark, the pipeline retained all 41 studies included in the original review; after audit adjustment, recall was 100.0% (46/46), accuracy 99.4% (156/157), precision 97.9% (46/47), and F1 98.9%. For clozapine/suicidality, recall was 79.3% (46/58), and F1 was 76.0%, with missed studies largely attributable to missing or non-informative abstracts. For clozapine patient/caregiver perspectives, recall was 88.9% (56/63), and F1 was 83.6%, with similar abstract-level constraints. Abstract-level extraction recovered audited metadata fields without detected errors and generated evidence maps that were thematically concordant with the main narrative structure of the reference reviews. Conclusions As a proof-of-concept, a fully local LLM pipeline can support scalable and auditable abstract-based scoping and high-level evidence mapping. Because performance was benchmarked against three reviews with partly audit-adjusted reference sets, the findings require confirmation in larger, independently adjudicated evaluations. Random human audit remains advisable, and expert full-text synthesis remains necessary when abstracts are non-informative or when mechanistic precision is required.

PMID:
42483140
Bibliographic data and abstract were imported from PubMed on 22 Jul 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 5
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement