Authors
Christopher Wong, Divy Kumar, Amiin Muse, Philip Wang, Finn Wintz, Sasha-Ann East, Jason Lyou, Naveena Yanamala, Zach Wood-Doughty
Published in
JAMIA open. Volume 9. Issue 5. Pages ooag143. Epub Oct 06, 2026.
Abstract
To simultaneously evaluate and compare LLM performance in clinical reasoning with a minimally necessary criteria for clinical alignment with patients.
We adapt the MedNLI dataset to create MedNLI-Pain, a new benchmark task that evaluates if LLMs that make accurate clinical inferences about patients can also anticipate when those patients are in pain. MedNLI-Pain consists of patient vignettes originating from MedNLI that are then annotated with physician assessments of the described patient's pain.
All LLMs had higher agreement with physicians in MedNLI than in MedNLI-Pain and its counterfactually augmented subset. This trend persisted when accounting for the higher difficulty of MedNLI-Pain by normalizing to human performance. Large language model performance improved with more training data in MedNLI but not MedNLI-Pain.
Large language models may lack necessary capabilities (eg, recognizing patients' pain) for clinical alignment. Success at a clinical reasoning task, in this case medical natural language inference, may not entail commensurate ability to align with patients' interests.
Further research on alignment with patients is needed because the implications may direct future research into either closed-ended or agentic applications.
PMID:
42840750
Bibliographic data and abstract were imported from PubMed on 07 Oct 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 8
- Comments 0