Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

A Cautious Integration With AI in the Clinic: A Standardized-Patient Pilot Study of ChatGPT's Reliability in Hamilton Depression Rating Scale Scoring.

Created on 04 Sep 2026

Authors

Chun-Hung Chang, Szu-Wei Cheng, Wei-Jen Chen, Chung-Wen Chang, Ting-Hui Liu, Jia-Hau Lee, Sheng-Che Lin, Kuan-Pin Su

Published in

Alpha psychiatry. Volume 27. Issue 4. Pages 51483. Epub Aug 13, 2026.

Abstract

Artificial intelligence (AI) integration offers significant potential to improve mental healthcare, however, the reliability of large language models (LLMs) in performing nuanced clinical tasks remains an important and largely unanswered question. This study aimed to evaluate ChatGPT's performance in scoring the Hamilton Depression Rating Scale (HAMD-21) compared with expert raters using standardized patients (SPs).
Three senior mental health experts created and portrayed scenarios for ten SPs representing diverse depressive symptom profiles. Recorded interviews were transcribed and used as input for ChatGPT-4o. HAMD-21 scores generated by ChatGPT were compared with those assigned by expert raters and with predefined script-based reference scores. Inter-rater reliability was assessed using intraclass correlation coefficient (ICC), and differences between raters were evaluated using Steiger's tests.
ChatGPT and the expert raters achieved good-to-excellent reliability for total HAMD-21 scores (experts: ICC = 0.9921; ChatGPT: ICC = 0.9739). However, expert raters achieved perfect ICCs on 11 individual items, whereas ChatGPT achieved perfect agreement on only 2 items. Steiger's test demonstrated that experts significantly outperformed ChatGPT on 10 individual items as well as on total scores (Z = 1.931, p = 0.0268). Qualitative review revealed that ChatGPT tended to overestimate scores on items related to insomnia and somatic symptoms (items 4-6 and 13) and frequently miscalculated total scores.
ChatGPT demonstrated excellent agreement on total HAMD-21 scores in structured, text-based depression assessments, supporting the potential role of LLMs as adjunctive tools for standardized depression severity evaluation. However, item-level discrepancies and systematic scoring errors indicate that human oversight remains essential for clinically nuanced interpretation.

PMID:
42694876
Bibliographic data and abstract were imported from PubMed on 04 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 4
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement