Authors
Edip Bayrak, Ramazan Azim Okyay, Hilmi Erdem Sumbul, Umida Khojakulova, Burhan Fatih Kocyigit
Published in
Rheumatology international. Volume 46. Issue 10. Oct 09, 2026. Epub Oct 09, 2026.
Abstract
This study compared the diagnostic accuracy, initial test selection, and initial treatment choices of three current-generation Large Language Models (LLMs) with those of two infectious disease specialists in managing infectious complications in patients with immunosuppressed rheumatic disease. This cross-sectional comparative study utilized 100 standardized Turkish-language clinical vignettes, developed by an expert committee, encompassing ten infection categories: viral, bone/joint, respiratory, central nervous system, gastrointestinal, genitourinary, skin and soft tissue, rare diseases, vaccination and prophylaxis, and adverse drug reactions. A prespecified reference standard, based on current clinical practice guidelines (ACR, EULAR, KLİMİK, IDSA, ATS, WHO, SANJO/EBJIS), was established prior to data collection. Three LLMs (Gemini 3.1 Pro, ChatGPT 5.5, Claude 4.7 Opus) and two blinded infectious disease specialists, operating under closed-book conditions, answered the same vignettes across three domains: diagnosis, initial diagnostic work-up, and initial treatment. Responses were evaluated dichotomously as correct or incorrect according to the reference standard. Diagnostic accuracy was 91% and 93% for the two human experts, compared to 99%, 97%, and 100% for Gemini 3.1 Pro, ChatGPT 5.5, and Claude 4.7 Opus, respectively. This pattern was consistent across all three domains (Cochran Q, all p ≤ 0.003). Human experts achieved complete case management in 83% and 90% of vignettes, while the LLMs achieved 97-100% (p < 0.001). Of 40 pairwise comparisons, 10 showed Bonferroni-corrected significance, all in human-versus-model comparisons. In this vignette-based comparison, all three LLMs answered a higher proportion of items correctly than the two participating infectious disease specialists in all three domains. As the cases were standardized written vignettes with a single prespecified correct answer, and only two human comparators were available, this reflects performance on a structured written task rather than superiority in clinical care. These findings suggest that LLMs could serve as a complementary decision-support tool, especially for identifying rare or opportunistic diagnoses, but should not replace specialists without further prospective, real-world validation.
PMID:
42853428
Bibliographic data and abstract were imported from PubMed on 10 Oct 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 5
- Comments 0