Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Bridging the Generalist-Subspecialist Gap with GPT-5-Thinking: Dual-Center Evaluation in Orbital and Head-and-Neck Tumor MRI Reports.

Created on 15 Sep 2026

Authors

Jie Li, Liang Xiao, Xiaoxia Qu, Lianze Du, Tingting Gong, Yao Yu, Yatong Li, Pengfei Xie, Guangyu Chu, He Li, Yidan Zhang, Hanxue Gao, Qinghai Yuan, Qinghe Han, Junfang Xian, Jianhua Liu

Published in

Radiology. Volume 320. Issue 3. Pages e253579.

Abstract

Background Identifying tumor histologic subtypes in complex regions is challenging for generalist radiologists. Whether reasoning large language models can bridge this expertise gap by interpreting radiologic descriptions remains underexplored. Purpose To evaluate whether Generative Pretrained Transformer (GPT)-5-Thinking can perform similarly to subspecialists and help bridge the generalist expertise gap in interpretation of MRI scans of orbital and head-and-neck tumors. Materials and Methods This dual-center retrospective study was performed at a generalist practice center and a subspecialist practice center and included convenience series of patients with pathologically confirmed orbital or head-and-neck tumors who had pretreatment MRI reports. GPT-5-Thinking processed narrative MRI reports to output benign-malignant classification and top-3 differential diagnoses. Phase 1 compared model accuracy with routine clinical report diagnoses in generalist and subspecialist practice settings. In Phase 2, seven generalists interpreted selected tumor reports twice, without and with GPT-5-Thinking assistance. McNemar tests and mixed-effects logistic regression models were used for analysis. Results This study included 1000 patients (mean age ± SD, 51 years ± 17.0; 520 men). In phase 1, the top-1 accuracy of GPT-5-Thinking was similar to that of subspecialists in patients with orbital tumors (76.7% [230 of 300] vs 78.3% [235 of 300]; P = .64) and head-and-neck tumors (67.5% [135 of 200] vs 68.0% [136 of 200]; P = .92). GPT-5-Thinking had higher top-1 accuracy than did routine generalists for patients with orbital tumors (74.0% [222 of 300] vs 68.7% [206 of 300]; P = .03) and head-and-neck tumors (64.0% [128 of 200] vs 27.0% [54 of 200]; P < .001). In phase 2, GPT-5-Thinking assistance was associated with an increase in generalist top-1 diagnostic accuracy from 61.4% (430 of 700) to 70.3% (492 of 700) (P < .001) in patients with orbital tumors and from 47.0% (329 of 700) to 61.0% (427 of 700) (P < .001) in patients with head-and-neck tumors. Conclusion Using MRI reports of patients with orbital or head-and-neck tumors, Generative Pretrained Transformer (GPT)-5-Thinking showed diagnostic accuracy similar to that of subspecialists; GPT-5-Thinking assistance was also associated with higher generalist diagnostic accuracy. © RSNA, 2026 Supplemental material is available for this article.

PMID:
42742382
Bibliographic data and abstract were imported from PubMed on 15 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 35
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement