Authors
Riccardo Lombardo, Maria Giovanna Stagliano, Marta Santioni, Vincenzo Pagliarulo, Giacomo Gallo, Silvia Secco, Paolo Dell'Oglio, Antonio Galfano, Alberto Olivero, Sabrina De Cillis, Alberto Quarà, Daniele Amparore, Francesco Porpiglia, Mattia Lo Re, Elettra Fuligni, Anna Cadenar, Andrea Cocci, Anton Zarraonandia, Leticia Ruibal Gago, Edoardo Sorba, Antonio Nacchia, Antonio Franco, Giorgio Ivan Russo, Simone Albisinni, Elisa De Lorenzis, Emanuele Montanari, Enrico Finazzi Agrò, Celeste Manfredi, Andrea Tubaro, Cosimo De Nunzio
Published in
International urology and nephrology. Sep 28, 2026. Epub Sep 28, 2026.
Abstract
This study aimed to develop and evaluate a ChatGPT-based application designed to assist multidisciplinary team (MDT) discussions in oncological urology.
A ChatGPT-based tool (UroMDT Advisor) was developed to: (1) generate case summaries, (2) provide evidence-based treatment suggestions aligned with EAU, ESMO, and NCCN guidelines, (3) offer supporting references, (4) compare opinions across specialties to identify inconsistencies, (5) draft final MDT reports, and (6) create patient summaries. A consecutive series of anonymized oncological cases from eight centers were entered into the system. For each case, GPT-generated opinions were compared with those of MDT participants (urologists, oncologists, radiotherapists) and the final consensus decision. The model operated in a blinded fashion during concordance testing and was subsequently unblinded to produce final reports and patient summaries. Two expert urologists independently assessed the accuracy, completeness, and clarity of GPT outputs.
A total of 180 cases were analyzed: 123 (68.3%) prostate, 21 (11.7%) renal, 21 (11.7%) urothelial, 12 (6.7%) testicular, and 3 (1.7%) penile cancer cases. Exact agreement between UroMDT Advisor and urologists, oncologists, radiation oncologists, and final MDT decisions was 61.7%, 66.7%, 63.3%, and 66.7%, respectively; the corresponding unweighted Cohen's kappa values were 0.401, 0.489, 0.441, and 0.482. After resolution of discordant ratings, the GPT-generated final reports were rated as accurate in 90% of cases, fairly accurate in 8%, and inaccurate in 2%. Completeness was rated high in 95%, moderate in 2%, and low in 3% of cases, while clarity was rated high in 95% and moderate in 5%.
The UroMDT Advisor showed moderate agreement with final MDT decisions and produced highly rated outputs. These findings support further evaluation of large language models as supervised decision-support and documentation tools in oncological MDT workflows while maintaining human oversight.
PMID:
42804130
Bibliographic data and abstract were imported from PubMed on 29 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 10
- Comments 0