Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Guideline-augmented prompting improves comparative preference and response consistency of large language model outputs for orthopaedic anaesthesia questions: A controlled prompting study.

Created on 04 Oct 2026

Authors

Anita Széll, Yinan Yu, Jacob F Oeding, Felix C Oettl, Adam Piasecki, Keti Dalla, Tobias Siöland, Mathias Hård Af Segerstad, Fredrik Olsen, Peter Larsson, Stefano Zaffagnini, Kristian Samuelsson, Fredrik Hessulf

Published in

Journal of experimental orthopaedics. Volume 13. Issue 4. Pages e70922. Epub Oct 03, 2026.

Abstract

Large language models (LLMs) are increasingly used in clinical contexts; however, performance in complex perioperative decision-making remains uncertain. Orthopaedic anaesthesia presents a demanding test case due to comorbidity burden and guideline-dependent management. Whether successive LLM generations and guideline-augmented prompting improve clinical alignment, comparative performance and response consistency was evaluated in this study.
In this controlled prompting study, 34 orthopaedic anaesthesia questions spanning seven clinical subdomains were independently answered by three LLMs (GPT-3.5-turbo, GPT-4o and GPT-5.2) and three expert anesthesiologists. GPT-4o and GPT-5.2 each generated responses with and without access to relevant clinical practice guidelines. All responses were anonymized and evaluated in blinded pairwise comparisons by an independent guideline-informed LLM-as-a-judge (GPT-5.2) using the same guideline material as the reference standard. The judge recorded preference, response consistency and confidence. Intra-model consistency was assessed from repeated independent generations of each question.
All LLMs were preferred over human experts in pairwise comparisons without guideline augmentation, with win rates of 67.6% (GPT-3.5-turbo), 79.4% (GPT-4o) and 95.1% (GPT-5.2). Performance was improved by guideline augmentation, most notably for GPT-4o (91.2%, +11.8 percentage points), while GPT-5.2 approached ceiling performance (97.1%). Response consistency varied between models. GPT-4o showed the highest baseline consistency (80.4%), whereas GPT-5.2 demonstrated no fully contradictory outputs but showed greater partial variability. Guideline augmentation numerically improved GPT-5.2 consistency (66.7% to 80.4%) and reduced inter-model differences. Directed qualitative analysis suggested that reviewer preference for later GPT generations was associated with greater completeness, explicit clinical reasoning and guideline-oriented responses.
Successive LLM generations demonstrated progressively improved performance in orthopaedic anaesthesia reasoning. Guideline-augmented prompting enhanced response quality, particularly for intermediate models. Guideline-informed LLM-as-a-judge evaluation appears promising for comparative assessment but requires further validation.
NA.

PMID:
42829621
Bibliographic data and abstract were imported from PubMed on 04 Oct 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 32
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement