Authors
Gökhan Çeker, Afonso Morgado, Giorgio Ivan Russo, Marco Falcone, Esther García Rojo, Celeste Manfredi, Omar Almidani, Rashed Rowaiee, EAU Young Academic Urologists (YAU) Sexual and Reproductive Health Working Group
Published in
International journal of impotence research. Aug 07, 2026. Epub Aug 07, 2026.
Abstract
Artificial intelligence (AI)-based language models are increasingly explored as tools for interpreting and applying clinical guideline recommendations. In urology, the European Association of Urology (EAU) recently introduced a guideline-specific chatbot; however, its comparative performance relative to contemporary general-purpose large language models (LLMs) remains unclear. In this structured comparative study, five AI systems-the EAU Guidelines Bot, ChatGPT-5, Gemini 2.5 Pro, Copilot - Smart GPT-5, and Perplexity Pro-were evaluated using 13 clinical questions derived directly from strongly recommended statements in the EAU erectile dysfunction (ED) guidelines. Responses were independently assessed by three senior reviewers across five predefined domains: relevance, clarity, structure, clinical utility, and factual accuracy, using a 5-point Likert scale. The primary outcome of the study was the composite performance score, which was calculated as the mean of the five domain scores. Inter-rater reliability was calculated using ICC(2,k), and differences among models were analyzed with the Friedman test followed by Holm-adjusted Wilcoxon post-hoc comparisons. Significant performance differences were observed across all domains (all p < 0.001). The highest composite scores were observed for Gemini 2.5 Pro [4.60 (4.40-4.73)] and the EAU Guidelines Bot [4.53 (4.47-4.80)], followed by ChatGPT-5 [4.27 (4.07-4.47)]. Lower composite scores were observed for Copilot - Smart GPT-5 [3.73 (3.40-3.87)] and Perplexity Pro [3.60 (3.47-3.80)]. Domain-level analysis showed consistently high median scores (≥ 4) for factual accuracy among top-performing models, whereas variability was more pronounced in clarity, structure, and clinical utility. These findings suggest that both guideline-specific systems and advanced general-purpose LLMs may generate responses broadly consistent with guideline-based recommendations in structured ED scenarios. However, variability across domains-particularly in structure and clinical utility-and modest differences in composite performance suggest that these models should be interpreted as supportive tools rather than definitive clinical decision-making systems, requiring further validation in real-world settings.
PMID:
42567933
Bibliographic data and abstract were imported from PubMed on 08 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 2
- Comments 0