Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Comparative Performance of Large Language Models in Urodynamic Trace Interpretation Using an Adapted Expert-Led Evaluation of Generative AI Competence and Excellence Framework.

Created on 05 Oct 2026

Authors

Oguzhan Akpinar, Anil Erkan, Alper Keskin, Muhammet Guzelsoy

Published in

Neurourology and urodynamics. Oct 05, 2026. Epub Oct 05, 2026.

Abstract

Interpretation of urodynamic studies is complex and subject to substantial inter-observer variability, even among experienced clinicians. While artificial intelligence has shown promise in urology, evidence regarding the performance of large language models (LLMs) in urodynamic trace interpretation remains limited. This study aimed to systematically compare the performance of multiple contemporary LLMs in interpreting standard urodynamic traces using an adapted ELEGANCE framework.
In this retrospective, single-centre comparative study, 119 anonymised urodynamic machine printouts retrieved from the institutional patient archive were independently interpreted by five LLM-based platforms (ChatGPT-5.2, Gemini 3, Perplexity AI Pro, DeepSeek-V3.2, and Grok 4.1). Case-specific expert reference interpretations were produced by one of three urologists, whereas all LLM-generated interpretations were independently evaluated by all three urologists. LLM-generated interpretations were assessed using an adapted version of the "Expert-Led Evaluation of Generative AI Competence and Excellence" (ELEGANCE) questionnaire, which evaluates relevance, completeness, applicability, structure, language/terminology, satisfaction, and hallucination. Between-model comparisons were performed using Friedman and post-hoc Wilcoxon signed-rank tests.
Overall ELEGANCE scores differed significantly among models (χ2[4] = 12.67, p = 0.013). Gemini achieved the highest mean score (25.77 ± 3.74), followed by Perplexity, DeepSeek, and ChatGPT-5.2, while Grok demonstrated significantly lower performance. Gemini consistently outperformed other models in clinically oriented domains, including relevance, completeness, and applicability (all p < 0.05). No hallucinations were identified in any model output. Inter-rater reliability among the three expert reviewers was good (average-measures ICC = 0.898, 95% CI 0.853-0.929; p < 0.001).
Contemporary LLMs can generate structured interpretations of urodynamic traces, with meaningful performance differences across the evaluated platforms. While strengths in structure and terminology are evident, limitations in clinical reasoning persist. LLMs may serve as valuable assistive tools for standardisation and quality assurance in urodynamics but are not yet suitable for autonomous clinical decision-making.

PMID:
42831263
Bibliographic data and abstract were imported from PubMed on 05 Oct 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 2
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement