Authors
Phillip D Jenkins, Steven Bedrick, Lisa Karstens, William Hersh, Bridge2AI-Voice Consortium, David A Dorr
Published in
Frontiers in digital health. Volume 8. Pages 1846369. Epub Jul 01, 2026.
Abstract
The human voice contains rich acoustic information indicative of laryngeal pathology, yet current screening relies on resource-intensive in-person laryngoscopy. While artificial intelligence has shown promise for voice analysis, progress has been limited by small, inconsistent datasets and challenges to clinical translation. The Bridge2AI-Voice initiative addresses these barriers by providing a large-scale, ethically sourced dataset with standardized, privacy-preserving derived features.
To determine whether the derived-feature release of Bridge2AI-Voice v3.0.0 can support a high-sensitivity screening model for laryngeal lesions and to evaluate its translational readiness using telemedicine implementation frameworks.
We analyzed data from 205 adult participants (136 controls, 52 benign vocal fold lesions, 13 precancerous lesions, 4 laryngeal cancer) drawn from the Bridge2AI-Voice v3.0.0 derived-feature release. An L2-regularized logistic regression model was fit to 131 OpenSMILE static acoustic features with age and sex at birth, evaluated under participant-level stratified 10-fold nested cross-validation. Inner-fold cross-validation was used for operating-point threshold selection. Pre-specified validity tests against age confounding included a DeLong comparison against an age-only baseline and an age-stratified label permutation test. Alternative feature modalities (SPARC articulatory features, Mel spectrogram derivatives, and multimodal combinations) and alternative classifier families were evaluated as robustness checks.
The OpenSMILE-based model achieved cross-validated AUC 0.812 (95% CI 0.744-0.876), with operating-point sensitivity 0.870 (95% CI 0.767-0.939) and specificity 0.566 (95% CI 0.479-0.651). Model discrimination significantly exceeded an age-only baseline (DeLong p = 0.0008) and survived age-stratified label permutation (observed AUC 0.812 vs. null mean 0.553, p = 0.0099). Subgroup analysis showed approximately consistent sensitivity across benign (0.865) and precancerous (0.846) lesion subgroups. Alternative feature modalities did not provide incremental discriminative information beyond OpenSMILE, and alternative classifier families produced AUCs within bootstrap confidence intervals of the primary model.
Derived acoustic features from the Bridge2AI-Voice v3.0.0 release combined with basic demographic information support cross-validated discrimination of vocal fold lesions consistent with the upper range of published voice-based laryngeal pathology classifiers. The result is presented as a candidate signal warranting confirmatory investigation in a larger, prospectively recruited cohort.
PMID:
42459998
Bibliographic data and abstract were imported from PubMed on 16 Jul 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 4
- Comments 0