Authors
Ángela Abejez-Arrizabalaga, Galadriel Pellejero-Sagastizabal, Rocío Aznar-Gimeno, Beatriz Franco, Santiago Letona-Giménez, Elena Morte-Romea, Julia Origüen, María Teresa Pérez-Rodríguez, Pilar Retamar, Esmeé Ruizendaal, Teske Schoffelen, Jeroen Schouten, Natalia Sisamón, Rafael Del Hoyo, José Ramón Paño-Pardo
Published in
Clinical microbiology and infection : the official publication of the European Society of Clinical Microbiology and Infectious Diseases. Aug 22, 2026. Epub Aug 22, 2026.
Abstract
Evaluate whether general-purpose large language models (LLMs) demonstrate competencies suitable for antimicrobial stewardship (AMS) support and characterize their failure modes.
Cross-sectional evaluation of seven LLMs (GPT-5, Claude Sonnet 4.5, Gemini 2.5 Pro, Grok 4, Llama-3.3-70b-instruct, Qwen 2.5-72b-instruct, DeepSeek-chat-v3.1) using 30 clinical scenarios mapped to ESCMID AMS competency frameworks. Scenarios included deliberate traps for fabrication and dangerous recommendations. Six AMS experts from the Netherlands and Spain performed blinded dual evaluation using content scores (0-5 scale) and binary safety flags for fabrication and danger. Standard and incentivizing prompt framings were compared.
Four commercial models achieved mean content scores above 3.9/5.0: Claude Sonnet 4.5 (4.06), Gemini 2.5 Pro (3.96), Grok 4 (3.96), and GPT-5 (3.94). Open-weight models scored significantly lower (2.94-3.57). No model achieved more than 63% responses free of fabrication or danger flags. However, fabrication did not impair clinical utility in non-trap scenarios (all within-category comparisons p>0.20). Danger flags ranged from 6.7% to 16.7% across models, with no significant difference between commercial and open-weight models. Incentivizing prompts were associated with a consistent 0.48-point-content score improvement (p=0.006), though significance attenuated after accounting for scenario-level clustering. Evaluators endorsed LLMs as useful AMS support tools with moderate supervision (5/6), identifying documentation preparation and trainee education as promising applications.
Medically untrained LLMs demonstrate competencies suitable for supervised AMS support. Fabrication remains the central safety challenge and requires verification workflows; danger, though less frequent (6.7-16.7%), concentrated in identifiable and therefore mitigable failure modes. Non-clinical stewardship tasks (education, documentation, communication) can benefit now, whereas clinical recommendations require expert oversight. Mapping these boundaries allows AMS teams, particularly those understaffed or without on-site infectious diseases expertise, to decide where LLM support adds value rather than risk.
PMID:
42632419
Bibliographic data and abstract were imported from PubMed on 23 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 6
- Comments 0