Authors
Zixuan Yu, Jacqueline Chipkin, Julia Siar, Suelynn Emily Ren, Fei Xia, Meliha Yetisgen, Nicholas J Dobbins, Taylor M Black
Published in
Canadian journal of psychiatry. Revue canadienne de psychiatrie. Pages 7067437261479118. Aug 31, 2026. Epub Aug 31, 2026.
Abstract
BackgroundWe evaluated whether summaries and large language models (LLMs) preserve predictive performance for inpatient violence risk and assessed performance across demographics.MethodsWe conducted a retrospective study of inpatient encounters at Harborview Medical Center between March 2021 and September 2023. Cases were inpatient behavioural violence events, and each was matched to 10 control encounters. For each encounter, we used clinical notes from the 72 hours before the event for cases and a matched window for controls. We compared Clinical-Longformer trained on original notes with pipelines that first generated summaries and then classified risk. Summaries were either general or guided by clinician-defined risk entities. We also tested Llama-3.1-8B and MedGemma-27B in zero-shot (ZS) and fine-tuned settings. Primary endpoints were positive-class precision, recall, and F1. We also reported overall and subgroup area under the receiver operating characteristic curve (AUROC) with bootstrap confidence intervals for sex, race, ethnicity, age, and mental health flag.ResultsClinical-Longformer on original notes achieved the strongest performance (F1 = 0.776, AUROC = 0.959). Replacing original notes with general or entity-guided summaries reduced performance (F1 = 0.465 and 0.510). Llama-3.1-8B improved with sequence-classification fine-tuning on original notes but remained below Clinical-Longformer (F1 = 0.701; AUROC = 0.933), while ZS prompting performed poorly. MedGemma-27B also remained inferior, with its best performance from sequence-classification fine-tuning on original notes (F1 = 0.612, AUROC = 0.947). Subgroup AUROCs were high across sex, race, ethnicity, and age, with lower performance among patients with a documented mental health flag.ConclusionDirect classification of original notes remained the most reliable strategy. Summaries of clinical notes may compress or alter cues needed for discrimination, and off-the-shelf prompting was inadequate as a stand-alone predictor. A practical path is to anchor risk scoring in long-context discriminative models and use LLMs for auditable, clinician-facing summaries or rationales. Prospective and external validation are needed before clinical deployment.
PMID:
42671866
Bibliographic data and abstract were imported from PubMed on 01 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 17
- Comments 0