Authors
Kota Murakami, Ryosuke Matsuzawa, Yuya Tahara-Arai, Haruka Ozaki
Published in
JMIR medical education. Volume 12. Pages e84266. Sep 11, 2026. Epub Sep 11, 2026.
Abstract
Cloud-hosted large vision-language models (LVLMs) often outperform open-weight models on multimodal benchmarks, but their applicability to health care examinations that require both text and image reasoning remains unclear. Existing studies on Japan's National Examination for Clinical Laboratory Technicians mainly used text-only settings and a limited set of models.
This study aimed to evaluate the accuracy and applicability of cloud-hosted LVLMs on this examination, quantify the contribution of images, examine alignment with human performance across medical subfields, and test generalizability on questions administered after the models' knowledge cutoff dates. We aim to provide a standardized benchmark of baseline multimodal performance under deployment-realistic conditions without additional fine-tuning.
For this comparative benchmarking study, we built the Kensagishi question-answer dataset (Kensagishi QA), a benchmark of 800 questions from the examination (2020-2023), including text-only and image-based questions, in a standardized structured JSON format, and strict scoring criteria (all correct options required). We tested GPT-5, GPT-4o, GPT-4o-mini, Gemini 1.5 Pro, Gemini 2.5 Pro, and Neva-22B using a uniform "numbers-only" response prompt; images were provided as Base64-encoded data when available. Accuracy was computed overall, by item type, and by medical subfield. Human reference was drawn from published examination results. Spearman correlation assessed model-human alignment. Generalization was evaluated on 200 questions from the 71st examination (2025), released after the models' knowledge cutoffs.
Top-performing models exceeded the 60% passing threshold. GPT-5 achieved 93.9% (95% CI 92.0%-95.4%) accuracy with images and 88.2% (95% CI 85.7%-90.2%) without images, Gemini 2.5 Pro 92.5% (95% CI 90.5%-94.1%) and 86.8% (95% CI 84.2%-88.9%), and GPT-4o 81.7% (95% CI 78.8%-84.2%) and 79.8% (95% CI 76.8%-82.4%), respectively. Providing images markedly improved accuracy for GPT-5, Gemini 2.5 Pro, and GPT-4o, whereas Neva-22B showed no improvement. Model-human rank correlations by medical subfield were moderate for GPT-4o (Spearman ρ=0.61; 95% CI -0.03 to 0.90; P=.06 with images) and statistically significant for Gemini 1.5 Pro without images (ρ=0.67; 95% CI 0.07-0.91; P=.03). On 200 postknowledge cutoff questions from the 71st examination (2025), model accuracies were comparable to historical averages (eg, GPT-5 91.5%, 95% CI 86.8%-94.6%; Gemini 2.5 Pro 90.7%, 95% CI 85.8%-94.0%; GPT-4o 82.0%, 95% CI 76.1%-86.7%), indicating no material performance degradation on unseen questions.
Cloud-hosted LVLMs show strong performance on this examination, and visual input is a key driver of accuracy for leading models. Models that better mirror human difficulty patterns (eg, GPT-4o) offer complementary insights for educational analytics, whereas the highest-accuracy models (eg, GPT-5 and Gemini 2.5 Pro) may be preferable when correctness is paramount. To our knowledge, this is the first systematic, multimodel evaluation isolating visual input that extends prior text-only, few-model studies and provides Kensagishi QA to guide educational deployment of medical AI.
PMID:
42779303
Bibliographic data and abstract were imported from PubMed on 24 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 3
- Comments 0