Authors
Kevin Li, Eric Sid, Qian Zhu
Published in
medRxiv : the preprint server for health sciences. Sep 20, 2026. Epub Sep 20, 2026.
Abstract
Rare diseases affect an estimated 300 million people worldwide, yet the research needed to guide diagnosis and treatment is often fragmented across multiple unstructured literature sources. Natural history studies (NHS) are a key source of this evidence, but manually extracting structured information from NHS publications can be tedious and does not scale.
We have developed a proof-of-concept for an information extraction pipeline testing three opensource large-language models (LLMs) -- Athena-v3-AWQ, Google's Gemma3-27B, and Meta's Llama-3.1-70B-Instruct, to extract key NHS characteristics from PubMed abstracts curated from a Chan Zuckerberg Initiative disease research state model corpus (302 gold-standard and 8,338 full-corpus abstracts), and compared the models on efficiency, extraction completeness, and expert-rated accuracy.
All three models processed abstracts with success rates exceeding 99%. However, Gemma achieved the best overall performance, with the highest expert-rated accuracy (68.0% of outputs rated "good" vs. 36.0% for Llama and 10.0% for Athena) and the fastest runtime on the full corpus (~16 minutes for 3,547 abstracts), despite Llama scoring higher on the automated Token F1 metric (0.874 vs. 0.723), highlighting a divergence between automated and human evaluation. Athena's lower performance was largely attributable to verbatim copying rather than synthesis of extracted content.
These findings illustrate how locally deployed open-source LLMs can extract structured NHS characteristics at scale, thus supporting their use to accelerate evidence synthesis in rare disease research.
PMID:
42779962
Bibliographic data and abstract were imported from PubMed on 24 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 4
- Comments 0