Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Measuring Consistency Between Large Language Models' Responses to Preventive Care Queries and Official US Preventive Services Task Force (USPSTF) Recommendations: Systematic Test Involving All USPSTF Preventive Care Topics via Simulated User Prompts.

Created on 30 Jul 2026

Authors

Tim Johnson, Wolfgang Gaissmaier

Published in

JMIR AI. Volume 5. Pages e87034. Jul 30, 2026. Epub Jul 30, 2026.

Abstract

Large language models (LLMs) have the potential to provide individualized preventive care guidance at scale. Research, however, has found mixed performance among a small set of LLMs queried about select preventive care activities. These findings call for testing a larger set of LLMs on a wider range of preventive care topics.
This study aims to assess whether various popular LLMs generate outputs about preventive care consistent with a comprehensive set of recommendations from the US Preventive Services Task Force (USPSTF).
We investigated whether 35 popular LLMs produced outputs consistent with all publicly available USPSTF recommendations (n=142) published as of May 2025. The study occurred in 2 waves (wave 1, 2025: 28 LLMs; wave 2, 2026: 10 LLMs; 3 LLMs overlapping across waves). LLMs received queries from simulated users who, in baseline prompts, sought nonbinding, hypothetical advice about whether to participate in particular preventive care activities given their inclusion in a relevant population. LLM raters assessed LLM-USPSTF concordance (interrater reliability, wave 1: κ=0.8893; wave 2: κ=0.9366). Wave 2 tested chain-of-thought, few-shot, and role-based prompts (3 variants each for 426 tests per prompting approach per model). Wave 2 also tested an iterative prompt that sought clarification about previous LLM responses and a prompt that eliminated the user's reference to nonbinding, hypothetical advice. Automated methods classified responses to detect sources of LLM-USPSTF discrepancy. Further tests prompted LLMs to rate preventive care activities for relevant populations using the USPSTF grading scale.
Focusing on cases where LLM raters agreed, the study found in its baseline prompts that the LLM with the highest concordance rate generated responses consistent with USPSTF recommendations in 66.92% (89/133) of tests in wave 1 and 87.77% (122/139) of tests in wave 2; the LLM with the lowest rate accorded with USPSTF recommendations in 45.19% (61/135) of tests in wave 1 and in 50.36% (69/137) of tests in wave 2. Eliminating reference to nonbinding, hypothetical advice did not alter the highest-performing model's rate of concordance (122/139, 87.77%). The highest concordance rate increased with chain-of-thought (404/421, 95.96%), role-based (371/416, 89%), and iterative prompting (127/140, 90.71%); however, it moderately decreased with few-shot prompting (361/419, 86.15%). Automated content analysis found high rates of LLMs avoiding definitive recommendations. When prompted to grade preventive care activities, the highest-performing LLM matched USPSTF grades in 85.92% (122/142) of tests in wave 1 and in 94.37% (134/142) of tests in wave 2.
LLMs' consistency with USPSTF recommendations varies. Deviations result mainly from LLMs' avoidance of definitive statements. LLM-USPSTF concordance has improved markedly in newer models, and this concordance increases with particular prompting approaches.

PMID:
42530970
Bibliographic data and abstract were imported from PubMed on 30 Jul 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 9
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement