Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Assessing Artificial Intelligence Language Models for Patient-Oriented Information on Chiari Malformation Types: A Structured Multinational Evaluation Based on Real-World Expectations.

Created on 22 Sep 2026

Authors

Mert Çetin, Ali Çağlar Turgut, Orhan Beger, Fang Wang, Aaron S Dumont, Fuyou Guo, Mehmet Turgut, R Shane Tubbs

Published in

World neurosurgery. Pages 125350. Sep 21, 2026. Epub Sep 21, 2026.

Abstract

To evaluate the performance of contemporary large language model (LLM)-based artificial intelligence (AI) systems in providing patient-oriented information across multiple Chiari malformation (CM) subtypes.
Five AI models (ChatGPT-4o, Gemini, Copilot, DeepSeek, and Perplexity AI) were evaluated using seven patient-oriented questions on the definition, epidemiology, etiology, symptomatology, diagnosis, treatment, and prognosis of CM. Nine CM subtypes were included: Types 0, 0.5, 1, 1.5, 2, 3, 3.5, 4, and 5. A total of 315 AI-generated responses were independently assessed by three blinded neurosurgeons using a three-point ordinal scoring system for accuracy, comprehensiveness, and conciseness.
There were significant differences among the AI models across all domains evaluated (all p<0.001). Gemini performed most strongly in accuracy and comprehensiveness, while Perplexity performed best in conciseness. Copilot generally performed worse across the domains evaluated. Significant subtype-based differences were also identified, with AI-generated responses concerning Types 1 and 2 performing substantially better than those concerning rarer subtypes. Diagnostic questions achieved the highest overall performance, while prognosis- and definition-related questions performed comparatively less well. Reviewer-based analyses revealed substantial variability in scoring behavior.
Contemporary LLMs perform variably in providing patient-oriented information about CMs. Although some AI systems generated relatively accurate and comprehensive responses for well-established CM subtypes, performance was lower for rarer and less clearly defined variants, highlighting the continuing importance of expert oversight in delivering AI-assisted medical information.

PMID:
42767556
Bibliographic data and abstract were imported from PubMed on 22 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 10
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement