Authors
Jinwan Shi, Yong Ma, Yinhui Liu, Bowen Shi
Published in
Frontiers in public health. Volume 14. Pages 1904062. Epub Aug 05, 2026.
Abstract
Large language models (LLMs) are entering drug-safety and regulatory workflows, yet their behavior at the boundary between unavailable evidence and risk reassurance remains poorly characterized. Drug-induced liver injury (DILI) is a stringent setting because evidence is fragmented across labels, case reports, mechanistic studies, and curated knowledge bases, while unsupported low-risk reassurance can be consequential.
We evaluated five LLMs on a fixed 47-drug DILI risk-assessment panel using 1,410 parsed responses from closed-book answering and an evidence-gated protocol that restricted responses to supplied PubMed-derived evidence. Curated DILI resources were excluded from prompts and used only for evaluation. After filtering, 38 drugs had no direct DILI decision-support evidence in the supplied packet, and 9 had direct DILI-relevant evidence; a post hoc PubMed title/abstract recall stress test identified five additional drugs with recoverable direct DILI evidence outside the packet.
In the no-direct-evidence slice defined by the supplied packet, closed-book models rarely abstained, with drug-level abstention ranging from 5.3% to 33.3%; the evidence-gated protocol required abstention, which all models followed for every no-direct-evidence drug. The same pattern held for recent or low-recognition drugs, where evidence-gated abstention reached 92.0% to 100.0% vs. 8.0% to 49.3% under closed-book answering. Closed-book models also produced high-confidence low-risk responses for DILI-positive drugs, a label-discordant pattern largely removed by evidence gating. Independent expert review of selected responses showed that label discordance did not always imply a clinically unreasonable low-risk category, but identified unsafe reassurance through overconfident wording and under-cautious responses in selected cases. When direct DILI evidence was provided, all models preserved citation-grounded non-abstaining answers. However, they differed in how often they committed to a conclusive rather than an uncertain risk category. Citation-bearing evidence-gated responses cited only the supplied PubMed identifiers and achieved 91.2% to 100.0% concordance with the supplied grade.
These findings identify unsupported reassurance as measurable evidence-boundary behavior in LLM drug-risk assessment and establish a reproducible framework for auditing adherence to an externally supplied evidence boundary, defined by PubMed evidence classification and enforced by prompt policy rather than inferred independently by the model.
PMID:
42620926
Bibliographic data and abstract were imported from PubMed on 20 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 1
- Comments 0