Authors
Emanuele Perrone, Giuseppe Parisi, Maria Consiglia Giuliano, Ilaria Capasso, Nicola Macellari, Luca Russo, Angela Santoro, Francesco Fanfani
Published in
JCO clinical cancer informatics. Volume 10. Issue 3. Pages e2600142. Epub Sep 16, 2026.
Abstract
To compare the concordance of ChatGPT, Gemini, and Claude with a prespecified expert guideline-based reference standard in fabricated endometrial cancer clinical vignettes under standardized prompting.
We conducted a case-based in silico benchmarking study using 35 fabricated postoperative endometrial cancer vignettes representing a broad spectrum of ESGO-ESTRO-ESP 2025 management scenarios. Each vignette was submitted to ChatGPT, Gemini, and Claude in independent chat sessions using the same standardized prompt. The primary end point was concordance with the prespecified expert reference standard, scored as 0 (discordant), 1 (partially concordant), or 2 (fully concordant). Secondary end points were major safety issues and recognition of missing critical information. A post hoc subgroup analysis evaluated the effect of guideline-informed prompting in 12 cases.
Concordance differed significantly across models (P < .001). Gemini achieved the highest performance, with a median concordance score of 2 (IQR 1-2), compared with 1 (IQR 0-1) for ChatGPT and 0 (IQR 0-1) for Claude. Fully concordant recommendations were generated in 65.7% of cases by Gemini, 2.9% by ChatGPT, and 8.6% by Claude. Major safety issues also differed across models (P = .007), occurring in 22.9% of Gemini responses, 48.6% of ChatGPT responses, and 54.3% of Claude responses. In the subset of vignettes with intentionally missing decisive information, Gemini identified the need for additional data in 83.3% of cases, compared with 50.0% for both ChatGPT and Claude. In the post hoc subgroup analysis, guideline-informed prompting significantly improved concordance for Gemini and Claude.
Mainstream consumer large language models showed substantially different performance in postoperative endometrial cancer decision making. Although Gemini achieved higher concordance and fewer major safety issues than ChatGPT and Claude, no model demonstrated performance sufficient to support autonomous clinical use in multidisciplinary management.
PMID:
42748382
Bibliographic data and abstract were imported from PubMed on 17 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 8
- Comments 0