Authors
Ryan K McBain, Ellice Huang, Caroline Figueroa, Li Ang Zhang, Jonathan Cantor
Published in
BMJ mental health. Volume 29. Issue 1. Sep 28, 2026. Epub Sep 28, 2026.
Abstract
Medical artificial intelligence (AI) benchmarks are increasingly used to assess the readiness of large language models for health-related tasks, but aggregate performance scores can obscure clinically meaningful variation across domains. HealthBench, an open benchmark of 5000 multi-turn health conversations, represents a salient example: the score is difficult to interpret for decisions about specific clinical use cases, including those with mental health needs. Conversations involving suicidality, psychosis, eating disorders, trauma or relational dependence require recognising ambiguous disclosures, responding safely under uncertainty, avoiding reinforcement of harmful beliefs or over-reliance and knowing when to escalate to human-based crisis support. We propose that HealthBench and similar benchmarks report domain-specific results. To demonstrate feasibility and utility, we adopted OpenAI's classified dataset to identify 332 mental health conversations, validated against human review (sensitivity 0.980, specificity 1.000, positive predictive value 1.000, negative predictive value 0.980, F1 0.990, inter-rater agreement 99.0%). We analysed mental health domain performance separately from overall HealthBench performance using a stratified random sample of 500 conversations. Across deployments, mental health point estimates differed from overall estimates by -0.034 to +0.047. Mental health domain composition was also uneven: for example, postpartum depression represented 26.8% of mental health conversations compared with 4.5% for suicidal crisis, 0.9% for psychosis and 1.5% for eating disorders. The findings suggest benchmark averages can mask domain-specific insights relevant to clinical safety. We argue that interpretable medical AI evaluation requires transparent domain composition; subgroup and failure mode reporting; calibration against clinically meaningful anchors, including human expert performance; and reproducible versioned analyses as models, grader systems and deployment contexts evolve.
PMID:
42805769
Bibliographic data and abstract were imported from PubMed on 29 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 3
- Comments 0