Authors
Anita Széll, Yinan Yu, Jacob F Oeding, Felix C Oettl, Adam Piasecki, Keti Dalla, Tobias Siöland, Mathias Hård Af Segerstad, Fredrik Olsen, Peter Larsson, Stefano Zaffagnini, Kristian Samuelsson, Fredrik Hessulf
Published in
Journal of experimental orthopaedics. Volume 13. Issue 4. Pages e70922. Epub Oct 03, 2026.
Abstract
Large language models (LLMs) are increasingly used in clinical contexts; however, performance in complex perioperative decision-making remains uncertain. Orthopaedic anaesthesia presents a demanding test case due to comorbidity burden and guideline-dependent management. Whether successive LLM generations and guideline-augmented prompting improve clinical alignment, comparative performance and response consistency was evaluated in this study.
In this controlled prompting study, 34 orthopaedic anaesthesia questions spanning seven clinical subdomains were independently answered by three LLMs (GPT-3.5-turbo, GPT-4o and GPT-5.2) and three expert anesthesiologists. GPT-4o and GPT-5.2 each generated responses with and without access to relevant clinical practice guidelines. All responses were anonymized and evaluated in blinded pairwise comparisons by an independent guideline-informed LLM-as-a-judge (GPT-5.2) using the same guideline material as the reference standard. The judge recorded preference, response consistency and confidence. Intra-model consistency was assessed from repeated independent generations of each question.
All LLMs were preferred over human experts in pairwise comparisons without guideline augmentation, with win rates of 67.6% (GPT-3.5-turbo), 79.4% (GPT-4o) and 95.1% (GPT-5.2). Performance was improved by guideline augmentation, most notably for GPT-4o (91.2%, +11.8 percentage points), while GPT-5.2 approached ceiling performance (97.1%). Response consistency varied between models. GPT-4o showed the highest baseline consistency (80.4%), whereas GPT-5.2 demonstrated no fully contradictory outputs but showed greater partial variability. Guideline augmentation numerically improved GPT-5.2 consistency (66.7% to 80.4%) and reduced inter-model differences. Directed qualitative analysis suggested that reviewer preference for later GPT generations was associated with greater completeness, explicit clinical reasoning and guideline-oriented responses.
Successive LLM generations demonstrated progressively improved performance in orthopaedic anaesthesia reasoning. Guideline-augmented prompting enhanced response quality, particularly for intermediate models. Guideline-informed LLM-as-a-judge evaluation appears promising for comparative assessment but requires further validation.
NA.
PMID:
42829621
Bibliographic data and abstract were imported from PubMed on 04 Oct 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 32
- Comments 0