Authors
Qingxia Wu, Qingxia Wu, Peipei Zhang, Zhifeng Yi, Yu Shen, Yan Bai, Hongna Tan, Pei Dong, Zhong Xue, Neil Roberts, Meiyun Wang
Published in
Journal of medical Internet research. Volume 28. Pages e92183. Aug 07, 2026. Epub Aug 07, 2026.
Abstract
Multimodal large language models are increasingly used in radiological diagnosis, but their performance has not been systematically evaluated across volumetric (3D) imaging, real-world clinical versus public teaching cases, and bilingual contexts.
The aim of the study is to develop a bilingual radiology benchmark and characterize the diagnostic performance of state-of-the-art multimodal large language models across input modality, clinical setting (public teaching vs routine clinical), disease rarity, and clinical-history language and to disentangle linguistic from clinical-content effects through a cross-linguistic control experiment.
We constructed RadM-Bench, comprising 720 cases evenly distributed across 9 radiological subspecialties: 360 English public teaching cases enriched in rare diseases (RadEdu) and 360 Chinese routine clinical cases (RealClin). In total, 4 proprietary models (GPT-4o, O3, Gemini-2-Flash, and Gemini-2.5-Flash-Thinking) and 6 open-source models (Qwen2.5-VL-72B/7B, InternVL3-78B/8B, Llama-4-Scout-17B-16E, and MedGemma-4B) were evaluated under 4 input conditions: clinical history alone, history with radiologist-selected 2D key images, and history with volumetric data sampled at 2 and 10 frames per second (fps). Each response was scored on a 4-tier 0-3 diagnostic-quality rubric by 2 board-certified radiologists blinded to model identity. Mean scores with bias-corrected and accelerated bootstrap 95% CIs are reported. To disentangle language from clinical content, all 360 RealClin histories were translated into English and re-evaluated, with paired comparisons by Wilcoxon signed-rank tests and Benjamini-Hochberg false-discovery-rate correction.
Mean performance remained below 1.5 on the 0-3 scale for all 10 models on both datasets. Adding radiologist-selected 2D key images to clinical history improved performance in all 10 models (+19.8% to +139.2%). In RealClin at fps=10, all 8 evaluable models scored lower with volumetric input than with the 2D-image baseline (-5.3% to -31.4%); MedGemma-4B and Llama-4-Scout-17B-16E could only be evaluated at fps=2 due to context-window and graphics processing unit-memory constraints. At fps=2, a total of 8 out of 10 models declined (-6.8% to -28.6%), while Qwen2.5-VL-7B and InternVL3-8B showed marginal improvements (+2.4% and +2.1%). Cross-dataset transfer diverged by model category: proprietary models declined from RadEdu to RealClin (eg, O3 with images: 1.14 to 0.79), whereas Chinese-centric open-source models improved (eg, InternVL3-78B: 0.48 to 0.75). The rare-disease premium observed in 9 out of 10 models in RadEdu reversed in RealClin, where common-disease scores exceeded rare-disease scores in 7 out of 10 models under history-only input. Translating RealClin histories into English produced a numerical decrease in mean score for all 10 models, which were statistically significant in 9 out of 10 models after false discovery rate correction, excluding a Chinese-language penalty.
Within the scope of this benchmark, multimodal inputs improved performance over clinical history alone, but performance gaps remain in volumetric data processing and cross-context generalization, with mean diagnostic performance across the 10 evaluated models remaining below clinically actionable levels on both datasets.
PMID:
42566748
Bibliographic data and abstract were imported from PubMed on 08 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 15
- Comments 0