Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Accuracy, Reliability, and Bloom's Taxonomy Performance of Seven Large Language Models on Microbiology Questions.

Created on 10 Aug 2026

Authors

Volodymyr Dvornyk, Olena Bolgova, Volodymyr Mavrych

Published in

Advances in medical education and practice. Volume 17. Pages 621664. Epub Aug 04, 2026.

Abstract

Large language models (LLMs) are increasingly used as learning resources in medical education, yet their performance and reliability in microbiology, a discipline with a broad, heterogeneous knowledge base, have not been systematically evaluated across multiple platforms.
To benchmark seven publicly available LLMs on microbiology multiple-choice questions (MCQs), assessing overall accuracy, test-retest reliability, topic-specific performance, and the relationship between cognitive complexity and model performance.
Seven LLMs (Claude 4.6 Sonnet, Gemini 3.0, ChatGPT-5.2, Grok 4, Copilot, DeepSeek V3, and Kimi K2) completed 200 MCQs distributed across 20 microbiology topics and five Bloom's taxonomy levels in three independent sessions separated by 24-hour intervals. A total of 4200 responses were analyzed. Statistical analysis included one-way ANOVA with Tukey's HSD post hoc tests, repeated-measures ANOVA, intraclass correlation coefficients (ICCs), and Pearson correlations.
The collective mean accuracy was 86.18%. Six of seven systems exceeded the 80% high-competency threshold; Claude (89.83%), Grok (89.50%), and GPT (88.83%) led the group. Gemini (72.33%) was the only underperforming system. Test-retest reliability varied dramatically: Claude achieved excellent ICC (0.966), while Gemini exhibited poor reliability (ICC = 0.290), with session-to-session fluctuations of up to 100 percentage points on individual topics. Microbial Cell (100%) was the easiest topic; Viral Genomics (61.9%) was the most challenging across all systems. A uniform decline at Bloom's Level 4 (Analyze) was observed across all LLMs, with no model exceeding 78%.
Contemporary LLMs demonstrate substantial knowledge of microbiology but differ markedly in reliability. Response consistency, alongside accuracy, should be a primary criterion for educational deployment. These findings are specific to microbiology MCQ performance and may not generalize to open-ended clinical reasoning.

PMID:
42572759
Bibliographic data and abstract were imported from PubMed on 10 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 6
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement