Authors
Qi Yong Hemis Ai, Hoi Ming Kwok, Ming-Yi Lu, Tracy T S Lau, Kuo Feng Hung, Lun M Wong, Tiffany Y So, Ann D King, Ho Sang Leung
Published in
Insights into imaging. Volume 17. Issue 1. Sep 21, 2026. Epub Sep 21, 2026.
Abstract
MRI is used for staging head and neck cancer (HNC), but assigning T- and N-category criteria requires specialised expertise, and so many institutions offer only descriptive MRI reports. This study assessed potentials of commercial large language models (LLMs) to stage oral cavity cancer (OCC) using descriptive MRI reports and compared their accuracy with human experts.
104 eligible MRI reports were processed by five commercial LLMs (ChatGPT5.4, ChatGPT5.0, ChatGPT4.1, Gemini3.1 and DeepSeekV3.2). T- and N-categories, and overall stage were extracted from outputs of the LLMs. MRI reports were also staged by two multidisciplinary team (MDT) members. Accuracies of the LLMs and MDT members for cancer staging were assessed against the reference staging standard and compared using the McNemar test.
The LLMs showed accuracy of 54.8-76.0% for T-categorisation, 63.5-88.5% for N-categorisation, and 53.8-80.8% for overall stage. Compared with DeepSeekV3.2, ChatGPT and Gemini3.1 showed significantly higher accuracy (p ≤ 0.001), except for Gemini3.1 for T-categorisation (p = 0.07). No differences in accuracy for staging between ChatGPT versions and between them and Gemini3.1 (p = 0.08 to > 0.99). MDT members showed accuracy of 74.0-76.9% for T-categorisation, 76.9-77.9% for N-categorisation and 68.3-71.2% for overall stage. The tested LLMs did not consistently outperform MDT members for staging.
Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. Additionally, Low MDT staging accuracy highlighted the need for structured reports that include clinical cancer staging to facilitate MDT assessment, thus potentially ensuring optimised disease management.
Question Can commercial LLMs accurately assign cancer stage based on descriptive MRI reports to overcome the lack of specialised expertise in clinical practice? Findings Tested LLMs demonstrated variable performances, ranging 54.8-76.0% for T-categorisation, 63.5-88.5% for N-categorisation and 53.8-80.8% for overall stage, and failed to consistently outperform MDT members. Critical relevance statement Variable tested LLM performance in staging OCC from descriptive MRI reports suggested that they are not suitable as standalone staging tools. MDT members' low performances highlight that structured MRI reports with clinical cancer staging are needed to improve multidisciplinary assessment.
PMID:
42768233
Bibliographic data and abstract were imported from PubMed on 22 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 9
- Comments 0