Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Clinical speech AI for schizophrenia: Diagnostic accuracy of speech- and language-based digital biomarkers in a PRISMA-DTA systematic review and meta-analysis.

Created on 04 Sep 2026

Authors

Xiaoqi Tang, Junmei Chen, Yuehe Huang, Qian Yao

Published in

International journal of medical informatics. Volume 222. Pages 106690. Aug 29, 2026. Epub Aug 29, 2026.

Abstract

Speech- and language-based digital biomarkers are increasingly proposed as scalable tools for mental health assessment, but their diagnostic accuracy and readiness for clinical translation in schizophrenia remain uncertain.
MEDLINE, Embase, PsycINFO, Web of Science, Scopus, IEEE Xplore, ACM Digital Library and citation searching were used without language or geographic restriction. Eligible full empirical journal articles and full conference proceedings evaluated automated speech/language classification in schizophrenia-spectrum disorders. After peer review, a second author independently re-screened random 10 % samples of title/abstract and full-text records and independently reassessed QUADAS-2. Patient-level 2 × 2 data were synthesized with a bivariate random-effects model; mixed comparators, segment-level observations and AUC-only datasets were kept separate. Additional exploratory sensitivity analyses conducted during revision addressed index-test risk of bias, patient-level splitting, model/metric selection and modality. Deeks funnel-plot asymmetry testing was performed for the primary set.
Forty-six reports representing 47 datasets were included. The primary patient-level healthy-control set (k = 15; median total N = 100, range 16-284) yielded sensitivity 0.804 (95 % CI 0.743-0.853), specificity 0.829 (0.800-0.855) and SROC AUC 0.865. Between-study heterogeneity was substantial for sensitivity (τ2 = 0.328) but limited for false-positive rate (τ2 = 0.021), with correlation ρ =  - 0.890 and prediction-region coordinate bounds of sensitivity 0.486-0.947 and false-positive rate 0.118-0.241 (specificity 0.759-0.882). After exclusion of one study with self-reported diagnoses, the exploratory psychiatric/mixed-comparator set (k = 3) yielded sensitivity 0.637 (0.505-0.751), specificity 0.882 (0.736-0.953) and SROC AUC 0.762; the random-effects correlation reached a boundary, so individual studies remain central to interpretation. Segment-level results (three studies; 940 segments) and all 12 AUC-only datasets were reported descriptively. No primary dataset provided verified independent external validation. Deeks testing did not indicate small-study asymmetry (t = 0.328, df = 13, P = 0.748).
Automated speech and language models show promising internal discrimination, particularly against healthy controls, but the evidence does not establish real-world diagnostic utility. Clinically credible evaluation now requires locked-model prospective external validation in diagnostic-uncertainty populations, calibration and decision-curve reporting, standardized multilingual/multidevice acquisition, fairness assessment, interpretable outputs and explicit human oversight.

PMID:
42691873
Bibliographic data and abstract were imported from PubMed on 04 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 5
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement