Authors
Takanobu Hirosawa, Yukinori Harada, Ren Kawamura, Taro Shimizu
Published in
JMIR formative research. Volume 10. Pages e98819. Sep 09, 2026. Epub Sep 09, 2026.
Abstract
Large language models (LLMs) are increasingly used to generate differential diagnoses from clinical narratives. However, LLM-based diagnostic clinical decision support systems still lack a quantitative measure of how strongly a diagnosis is supported by the available case description. Conditional perplexity score quantifies how predictable a target text is given in a preceding context, with lower scores indicating greater predictability. We hypothesized that this concept can be adapted to diagnostic reasoning by treating the prediagnostic case description as the context and a diagnosis as the target text.
This study aims to evaluate whether conditional perplexity scores, computed by an independent LLM and conditioned on case-report narratives, differ between physician-verified correct and incorrect LLM-generated diagnoses. Specifically, we hypothesized that the correct LLM-generated diagnosis verified by physicians would have lower conditional perplexity scores than incorrect LLM-generated differential diagnoses. A secondary outcome was to compare this scoring behavior across differential diagnosis lists generated by different LLMs.
We performed a preliminary computational analysis of 392 peer-reviewed diagnostic case reports published in the American Journal of Case Reports in 2022. For each case, the prediagnostic clinical description was used as the conditioning context, and the case report-defined final diagnoses were treated as the gold standard. Conditional perplexity scores for differential diagnosis lists previously generated by LLaMA2, Bard, and GPT-4 were computed using an independent longer-context LLM, Qwen2.5-1.5B. We compared case report-defined final diagnoses, correct LLM-generated diagnoses verified by physicians, and incorrect generated diagnoses using nonparametric comparisons and receiver operating characteristic analyses.
All 392 cases had complete case descriptions and case report-defined final diagnoses. Across the top-10 differential diagnosis lists generated by LLaMA2, Bard, and GPT-4, 823 correct LLM-generated diagnoses verified by physicians and 10,875 incorrect generated diagnoses were analyzed. Case report-defined final diagnoses had lower conditional perplexity scores than incorrect generated diagnoses (median 39.9, IQR 17.7-119.9 vs median 133.3, IQR 37.5-672.1). Correct LLM-generated diagnoses also had lower conditional perplexity scores than incorrect LLM-generated diagnoses (median 43.3, IQR 16.6-147.5 vs median 133.3, IQR 37.6-672.1). Candidate-level discrimination was moderate overall (area under the receiver operating characteristic curve [AUC] 0.666, 95% CI 0.644-0.689) and was the highest for GPT-4-generated differential diagnosis lists (AUC 0.678, 95% CI 0.652-0.705), followed by LLaMA2 (AUC 0.662, 95% CI 0.625-0.698) and Bard (AUC 0.648, 95% CI 0.617-0.681). In within-case analyses, correct diagnoses had lower conditional perplexity than the mean incorrect diagnosis in 88.1% (237/269) to 91.1% (195/214) of evaluable lists.
Conditional perplexity provided a moderate quantitative signal associated with physician-verified correctness but did not reliably rank the correct diagnosis ahead of the strongest incorrect candidate, limiting its use as a stand-alone reranking method.
PMID:
42715414
Bibliographic data and abstract were imported from PubMed on 10 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 1
- Comments 0