Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Performance and Hallucination Analysis of Large Language Models on European Anesthesiology Examinations: Cross-Sectional Comparative Study.

Created on 02 Sep 2026

Authors

Stefan Andrei, Thibault Giet, Alexis Belouard, Mihai Stefan, Mihai Popescu, Sébastien Tanaka, Philippe Montravers, Aurélie Gouel

Published in

JMIR formative research. Volume 10. Pages e95859. Sep 01, 2026. Epub Sep 01, 2026.

Abstract

Large language models (LLMs) have shown promising performance on medical examinations across specialties. However, comparative evaluations of current-generation LLMs across multiple European anesthesiology examinations, alongside structured assessment of hallucinations vs question-related confusion, remain lacking.
This study aimed to compare the performance of 4 state-of-the-art LLMs on anesthesiology and intensive medicine examination questions and assess their hallucination rates.
This computational comparative study analyzed 437 multiple-choice questions (1748 queries) from 3 sources: nurse anesthetist school examinations (infirmier anesthésiste diplômé d'État [registered nurse anesthetist]; n=100, 22.9%), European Diploma in Anaesthesiology and Intensive Care (EDAIC; n=219, 50.1%), and EDAIC On-Line Assessment (n=118, 27.0%). Each question was submitted to 4 LLMs (Claude Sonnet 4.5, Gemini 2.5 Pro, GPT-5, and Grok 4) using standardized prompts via default web interface settings. Responses were evaluated through structured consensus review by 2 examiners for accuracy, hallucinations, and question-related confusion. Statistical analysis included Friedman and Wilcoxon signed-rank tests with Holm-Bonferroni correction, the Cochran Q test, and generalized estimating equations.
Average success rates ranged from 86% (SD 18%) to 94% (SD 10%) across LLMs and examination types, exceeding the EDAIC part I passing threshold, representing substantial improvement over previously reported GPT-3.5 performance. For the EDAIC, overall intermodel differences were significant (Friedman χ23=13.9; P=.003; W=0.02), with Gemini outperforming GPT-5 as the only pairwise difference. Hallucination rates ranged from 11% (11/100) to 20.1% (44/219) without significant intermodel differences. All models exceeded the EDAIC passing threshold.
Current-generation LLMs demonstrated consistently high performance across multiple European anesthesiology examinations but continue to produce clinically relevant hallucinations, supporting their role as supervised educational tools rather than autonomous learning resources. These findings underscore the need for structured integration frameworks and systematic verification when deploying LLMs as learning tools in medical education.

PMID:
42678529
Bibliographic data and abstract were imported from PubMed on 02 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 7
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement