Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Diagnostic Performance of Large Language Models for Orthopedic-Related Rare Diseases and Their Impact on Physicians' Diagnostic Accuracy: 2-Stage Comparative Evaluation Study Based on the Chinese Rare Disease Catalog.

Created on 25 Jul 2026

Authors

Tusheng Li, Ziqian Ma, Baodong Wang, Ning Fan, Aobo Wang, Lei Zang

Published in

Journal of medical Internet research. Volume 28. Pages e92931. Jul 24, 2026. Epub Jul 24, 2026.

Abstract

Orthopedic-related rare diseases are difficult to diagnose because of their low prevalence, heterogeneous phenotypes, and fragmented knowledge. Large language models (LLMs) can serve as dynamic knowledge-support tools, but their diagnostic performance and effect on physicians' decision-making remain unclear.
This study aims to compare the diagnostic performance of advanced LLMs for orthopedic-related rare diseases and to evaluate the effect of a 2-stage LLM-assisted diagnostic workflow on physicians' diagnostic accuracy and subjective acceptance.
We selected 40 orthopedic-related rare diseases from the Chinese Rare Disease Catalog. A total of 4 general-purpose LLMs each generated 1 primary diagnosis and 5 differential diagnoses per case. Diagnostic accuracy, defined as a correct primary diagnosis, was compared using the Cochran Q test and pairwise McNemar tests with Bonferroni correction. A representative LLM was integrated into a 2-stage workflow involving 27 intermediate and 15 senior orthopedic physicians. Physicians first diagnosed all cases independently and then rediagnosed the same cases after reviewing nonauthoritative LLM suggestions. Physician diagnostic data were primarily analyzed using mixed-effects logistic regression at the individual-diagnosis level. Case-level group accuracy was additionally assessed using >50% and ≥2/3 accurate-physician thresholds. After both rounds, physicians completed an 8-item Likert-scale questionnaire assessing subjective acceptance and workflow perceptions.
Claude Sonnet 4.5, ChatGPT-5.0, and Gemini 2.5 Pro each achieved 90% (36/40) primary-diagnosis accuracy, whereas DeepSeek-V3.2 achieved 67.5% (27/40; Cochran Q P<.001). Before LLM assistance, mean physician-level accuracy was 42.22% for intermediate physicians and 58.67% for senior physicians; after assistance, it increased to 68.80% and 83.33%, respectively. In the primary mixed-effects logistic regression analysis, physician seniority group and LLM assistance stage were significantly associated with diagnostic correctness (both P<.001), whereas the group-by-stage interaction was not significant (P=.10). Secondary case-level analyses using the >50% threshold showed improvement from 40% (16/40) to 67.5% (27/40) for intermediate physicians and from 57.5% (23/40) to 82.5% (33/40) for senior physicians, with similar findings using the ≥2/3 threshold. Cases accurately diagnosed by all 3 agents increased from 16 to 27. The questionnaire showed high internal consistency (Cronbach α=0.902) and generally positive attitudes, with no significant differences between physician groups (P=.11 to P=.78).
LLMs achieved high diagnostic accuracy for orthopedic-related rare diseases. In the 2-stage LLM-assisted workflow, LLM assistance was associated with higher diagnostic correctness in both physician groups, although seniority-related differences in the magnitude of benefit require evaluation in larger studies. Senior physicians retained higher diagnostic correctness than intermediate physicians. Secondary case-level analyses suggested attenuation of group-level gaps in case-recognition patterns. Physicians reported broadly positive workflow perceptions. Given the same-day repeated-case design and potential short-term recall bias, these exploratory findings should be interpreted cautiously and warrant prospective randomized, crossover, washout-period, independent-case, or real-world evaluations of LLM-assisted diagnostic workflows in orthopedics.

PMID:
42497411
Bibliographic data and abstract were imported from PubMed on 25 Jul 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 5
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement