Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Application of machine learning techniques to automate speech mispronunciation detection in children and adolescents.

Created on 10 Sep 2026

Authors

Nazila Ameli, Mahshid Nik Ravesh, Luan Matheus Trindade Dalmazo, Karen Pollock, Daniel DeSantis, Manuel Lagravere, Hollis Lai

Published in

Journal of applied oral science : revista FOB. Volume 34. Pages e20260078. Epub Sep 07, 2026.

Abstract

Speech sound disorders (SSDs) are common in children and can affect academic and psychosocial outcomes. Clinical identification relies on expert auditory-perceptual assessment, which is time-intensive and may vary across raters. Automated screening tools could support triage when access to specialists is limited.
This study aimed to develop and evaluate a deep learning (DL) classifier for detecting mispronunciation errors in standardized pediatric speech recordings obtained using a structured fricative-focused word elicitation protocol, using expert-adjudicated labels as the reference standard.
In this cross-sectional study, we analyzed 100 participants (6-18 years) providing 1,800 standardized word recordings. Two expert speech-language pathologists (SLPs) labeled recordings as mispronunciation present vs. absent; disagreements were adjudicated by a third SLP. Inter- and intra-rater reliability were assessed. Audio was denoised and standardized to 16 kHz. A pretrained transformer speech model (WavLM Base+) was fine-tuned for recording-level binary classification. Class imbalance was addressed using augmentation, weighted sampling, and focal loss. Performance was assessed on a held-out test set at the recording level using accuracy, sensitivity, specificity, precision, F1-score, and area under the ROC curve (AUC).
The overall recording-level prevalence of speech sound errors was 16.15%, with the "th" category (/θ/ + /ð/) being the most frequently mispronounced phoneme group. No significant differences were observed by sex, age group (6-9, 10-13, 14-18 years), or malocclusion status (p > 0.05). Labeling showed strong agreement (Cohen's κ = 0.84; intra-rater reliability 0.89-0.92). On the test set (253 recordings), the model achieved 90.9% accuracy, 77.8% sensitivity, 98.2% specificity, 95.9% precision, and AUC = 0.936.
Fine-tuned pretrained speech representations demonstrated promising performance for screening pediatric mispronunciation errors under standardized recording conditions using expert labels. The high-specificity profile supports its use as a screening and triage decision-support tool, while future studies are needed to validate performance across broader clinical and real-world settings.

PMID:
42715461
Bibliographic data and abstract were imported from PubMed on 10 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 12
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement