Authors
Xiaoqi Tang, Junmei Chen, Yuehe Huang, Qian Yao
Published in
International journal of medical informatics. Volume 222. Pages 106690. Aug 29, 2026. Epub Aug 29, 2026.
Abstract
Speech- and language-based digital biomarkers are increasingly proposed as scalable tools for mental health assessment, but their diagnostic accuracy and readiness for clinical translation in schizophrenia remain uncertain.
MEDLINE, Embase, PsycINFO, Web of Science, Scopus, IEEE Xplore, ACM Digital Library and citation searching were used without language or geographic restriction. Eligible full empirical journal articles and full conference proceedings evaluated automated speech/language classification in schizophrenia-spectrum disorders. After peer review, a second author independently re-screened random 10 % samples of title/abstract and full-text records and independently reassessed QUADAS-2. Patient-level 2 × 2 data were synthesized with a bivariate random-effects model; mixed comparators, segment-level observations and AUC-only datasets were kept separate. Additional exploratory sensitivity analyses conducted during revision addressed index-test risk of bias, patient-level splitting, model/metric selection and modality. Deeks funnel-plot asymmetry testing was performed for the primary set.
Forty-six reports representing 47 datasets were included. The primary patient-level healthy-control set (k = 15; median total N = 100, range 16-284) yielded sensitivity 0.804 (95 % CI 0.743-0.853), specificity 0.829 (0.800-0.855) and SROC AUC 0.865. Between-study heterogeneity was substantial for sensitivity (τ2 = 0.328) but limited for false-positive rate (τ2 = 0.021), with correlation ρ = - 0.890 and prediction-region coordinate bounds of sensitivity 0.486-0.947 and false-positive rate 0.118-0.241 (specificity 0.759-0.882). After exclusion of one study with self-reported diagnoses, the exploratory psychiatric/mixed-comparator set (k = 3) yielded sensitivity 0.637 (0.505-0.751), specificity 0.882 (0.736-0.953) and SROC AUC 0.762; the random-effects correlation reached a boundary, so individual studies remain central to interpretation. Segment-level results (three studies; 940 segments) and all 12 AUC-only datasets were reported descriptively. No primary dataset provided verified independent external validation. Deeks testing did not indicate small-study asymmetry (t = 0.328, df = 13, P = 0.748).
Automated speech and language models show promising internal discrimination, particularly against healthy controls, but the evidence does not establish real-world diagnostic utility. Clinically credible evaluation now requires locked-model prospective external validation in diagnostic-uncertainty populations, calibration and decision-curve reporting, standardized multilingual/multidevice acquisition, fairness assessment, interpretable outputs and explicit human oversight.
PMID:
42691873
Bibliographic data and abstract were imported from PubMed on 04 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 5
- Comments 0