Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Advancing Evidence-Based Medicine for Population, Intervention, Comparison, and Outcome Element Recognition and Extraction in Medical Literature: Large Language Model Approach.

Created on 15 Aug 2026

Authors

Zeyuan Hao, Yifan Duan, Yu Wang

Published in

Journal of medical Internet research. Volume 28. Pages e91215. Aug 14, 2026. Epub Aug 14, 2026.

Abstract

The exponential expansion of biomedical literature has created an urgent need for efficient methods to recognize and extract population, intervention, comparison, and outcome (PICO) elements-the foundational elements of evidence-based medicine.
This study systematically evaluated 2 complementary approaches for automating PICO recognition and extraction in medical literature: prompt engineering optimization and parameter-efficient fine-tuning (PEFT) of large language models (LLMs).
We developed a dual-phase methodological framework: (1) systematic prompt optimization incorporating in-context learning, chain of thought (COT), and multipath reasoning strategies; and (2) PEFT of the LLM architecture using low-rank adaptation (LoRA), quantized LoRA, and freeze techniques. The PubMed-PICO and NICTA-PIBOSO benchmark datasets were used for recognition tasks, and the EBM-NLP dataset was used for extraction tasks. Performance metrics included precision, recall, and F1-score. F1-score was adopted as the major metric as it balances precision and recall.
For prompt engineering, COT achieved the overall best performance across both recognition and extraction tasks. For example, in the recognition task, COT obtained strong average F1-scores of 77.1% (SD 0.5%) for the population element and 84.5% (SD 0.4%) for the outcome element on PubMed-PICO. In the extraction task, COT achieved the highest average F1-score of 73.9% across 3 PICO elements (the population, intervention, and outcome elements) on EBM-NLP. These results suggest that, for smaller models such as those with 3B parameters, explicit step-by-step guidance in COT is more effective than more complex prompting strategies. In PEFT implementations, for example, LoRA achieved the best recognition performance (mean F1-score 91.7%, SD 0.3% for population) on PubMed-PICO, whereas quantized LoRA showed the best extraction capability (mean F1-score 79.3%, SD 0.5% for intervention) on EBM-NLP. Fine-tuned models achieved competitive performance across all datasets, with notable gains on NICTA-PIBOSO and EBM-NLP. PEFT further enhanced the model's overall performance compared with prompt engineering, with element-dependent differences across PICO categories.
Our findings indicate that LLMs can effectively automate PICO recognition and extraction through 2 complementary approaches. First, prompt engineering allows the model to perform tasks directly without altering its internal settings. Second, the PEFT method further unlocks their maximum performance potential by incorporating additional fine-tuning based on prompt engineering. This work makes significant advances and provides critical insights for optimizing methodological approaches in clinical applications related to or comprising PICO extraction and recognition tasks.

PMID:
42600130
Bibliographic data and abstract were imported from PubMed on 15 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 6
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement