Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Data Extraction From Oncology Imaging Reports by Large Language Models: A Comparative Accuracy Study.

Created on 11 Aug 2026

Authors

Lea P Passweg, Johannes M Schwenke, Christof M Schönenberger, Flavio Locher, Julia Picker, Manuel Dieterle, Benjamin Thiele, Dimitri Hasler, Alessia Danelli, Andreas M Schmitt, Tobias Heye, Thomas Stojanov, Matthias Briel, Benjamin Kasenda

Published in

JCO clinical cancer informatics. Volume 10. Issue 3. Pages e2600002. Epub Aug 10, 2026.

Abstract

Manual data extraction from clinical text is resource-intensive. Locally hosted large language models (LLMs) may offer a privacy-preserving solution, but their performance on non-English data remains unclear. We investigated whether the accuracy of locally hosted LLMs is noninferior to human accuracy when determining metastasis status and treatment response from German radiology reports.
In this retrospective comparative accuracy study, five locally hosted LLMs (llama3.3:70b, mistral-small:24b, qwq:32b, qwen3:32b, and gpt-oss:120b) were compared against humans. A ground truth was established via duplicate human extraction and adjudication of discrepancies by a senior oncologist. The study was conducted at a tertiary referral hospital in Switzerland. We randomly sampled 400 radiology reports from adult patients with cancer (computed tomography, magnetic resonance imaging, positron emission tomography) generated between January 2023 and May 2025 and split them into a prompt optimization set (n = 100) and test set (n = 300). Primary outcomes were noninferiority (5 percentage points [pp] margin) of LLM classification accuracy compared with human accuracy for metastasis status (presence/absence by anatomic site) and treatment response categories. Secondary outcomes included accuracy for primary tumor diagnosis and radiologic absence of tumor.
The analysis included 400 reports from 317 patients. In the test set (n = 300), the human accuracy for metastasis status was 98.4% (95% CI, 98.0 to 98.8). All LLMs were noninferior; gpt-oss:120b performed best (97.6% accuracy; difference, -0.8 pp [90% CI, -1.3 to -0.3 pp]). For response to treatment, the human accuracy was 86.0% (95% CI, 83.2 to 88.8). All LLMs were inferior; the most accurate model, gpt-oss:120b, achieved 78.3% (difference, -7.7 pp [90% CI, -11.6 to -3.8 pp]).
In this study, LLMs were noninferior to human accuracy for classification of metastasis status but were inferior for response to treatment assessment.

PMID:
42574694
Bibliographic data and abstract were imported from PubMed on 11 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 8
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement