Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Large language models for lumbar spondylolisthesis detection: a multi-center pilot comparative radiographic accuracy study.

Created on 13 Aug 2026

Authors

Rehan R Khan, Tony Tannoury, Rohith Ryali, Matheus Alves, Brian S Tao, Sameh Abolfotouh, Omar Alnori, Mohammad El-Sharkawi, Kamel Moufarrej, Margaret Li, Lila R Mollick, Nader El-Hajj, Daman P Dhunna, Matthew T Kim, Neil V Shah, Chadi Tannoury

Published in

European journal of orthopaedic surgery & traumatology : orthopedie traumatologie. Volume 36. Issue 1. Aug 13, 2026. Epub Aug 13, 2026.

Abstract

Lumbar spondylolisthesis remains a radiographic diagnosis. With the improvement of large language models (LLMs), such as ChatGPT, in interpreting multimodal images, many patients have turned to LLMs for diagnostic insight. Yet the diagnostic reliability of emerging LLMs for spinal pathologies remains understudied. This study sought to evaluate the performance of both ChatGPT-4o and ChatGPT-5 in detecting lumbar spondylolisthesis on standing lateral lumbar radiographs from a public imaging dataset evaluated by expert spine surgeon consensus.
From the VinDr-SpineXR dataset library, we extracted 200 standing lateral lumbar radiographs, including 100 labeled as spondylolisthesis-positive and 100 labeled as spondylolisthesis-negative. Five fellowship-trained spine surgeons independently reviewed all 200 radiographs. Expert consensus was defined as agreement by at least three surgeons. These same 200 radiographs were independently analyzed by GPT-4o and GPT-5 using a standardized binary prompt to assess presence or absence of spondylolisthesis. Diagnostic performance was assessed relative to expert surgeon consensus.
After spine surgeon consensus was established, spondylolisthesis was confirmed in 81% of VinDr-SpineXR spondylolisthesis-labeled positive radiographs, while 2% of VinDr-SpineXR spondylolisthesis-labeled negative radiographs were reclassified as positive. With expert surgeon consensus set as ground truth, ChatGPT-5 outperformed ChatGPT-4o evidenced by higher sensitivity (67.5% vs. 49.4%) and overall accuracy (61.5% vs. 55.5%), with minor difference in specificity (57.3% vs. 59.8%). Inter-rater agreement was higher with ChatGPT-5 (κ = 0.238) than GPT-4o (κ = 0.091).
ChatGPT-5 outperformed ChatGPT-4o in detecting lumbar spondylolisthesis. Yet, both LLMs remained limited compared to fellowship trained spine surgeons.

PMID:
42593567
Bibliographic data and abstract were imported from PubMed on 13 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 3
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement