Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Accuracy and Safety of Large Language Models in Endometrial Cancer Decision Making: A Case-Based In Silico Benchmarking Study.

Created on 17 Sep 2026

Authors

Emanuele Perrone, Giuseppe Parisi, Maria Consiglia Giuliano, Ilaria Capasso, Nicola Macellari, Luca Russo, Angela Santoro, Francesco Fanfani

Published in

JCO clinical cancer informatics. Volume 10. Issue 3. Pages e2600142. Epub Sep 16, 2026.

Abstract

To compare the concordance of ChatGPT, Gemini, and Claude with a prespecified expert guideline-based reference standard in fabricated endometrial cancer clinical vignettes under standardized prompting.
We conducted a case-based in silico benchmarking study using 35 fabricated postoperative endometrial cancer vignettes representing a broad spectrum of ESGO-ESTRO-ESP 2025 management scenarios. Each vignette was submitted to ChatGPT, Gemini, and Claude in independent chat sessions using the same standardized prompt. The primary end point was concordance with the prespecified expert reference standard, scored as 0 (discordant), 1 (partially concordant), or 2 (fully concordant). Secondary end points were major safety issues and recognition of missing critical information. A post hoc subgroup analysis evaluated the effect of guideline-informed prompting in 12 cases.
Concordance differed significantly across models (P < .001). Gemini achieved the highest performance, with a median concordance score of 2 (IQR 1-2), compared with 1 (IQR 0-1) for ChatGPT and 0 (IQR 0-1) for Claude. Fully concordant recommendations were generated in 65.7% of cases by Gemini, 2.9% by ChatGPT, and 8.6% by Claude. Major safety issues also differed across models (P = .007), occurring in 22.9% of Gemini responses, 48.6% of ChatGPT responses, and 54.3% of Claude responses. In the subset of vignettes with intentionally missing decisive information, Gemini identified the need for additional data in 83.3% of cases, compared with 50.0% for both ChatGPT and Claude. In the post hoc subgroup analysis, guideline-informed prompting significantly improved concordance for Gemini and Claude.
Mainstream consumer large language models showed substantially different performance in postoperative endometrial cancer decision making. Although Gemini achieved higher concordance and fewer major safety issues than ChatGPT and Claude, no model demonstrated performance sufficient to support autonomous clinical use in multidisciplinary management.

PMID:
42748382
Bibliographic data and abstract were imported from PubMed on 17 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 8
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement