Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Validation of FUDAN VOICE: Signal and Acoustic Feature Comparison with Computerized Speech Lab (CSL).

Created on 19 Sep 2026

Authors

Siyan Xu, Xia Hu, Rui Fang, Chen Chen, Jun Shao, Chunsheng Wei, Min Shu

Published in

Journal of voice : official journal of the Voice Foundation. Sep 18, 2026. Epub Sep 18, 2026.

Abstract

This study aimed to validate the consistency between audio acquired by a custom WeChat mini-program, FUDAN VOICE (named Voice Acquisition in our earlier study), and Computerized Speech Lab (CSL) recordings at both signal and acoustic feature levels, and to characterize aggregate signal deviations between the two recording pathways, encompassing contributions from hardware, platform-level processing, and environmental noise.
A simultaneous recording design was employed, capturing parallel audio from CSL and FUDAN VOICE across a soundproof room and a quiet office environment. Signal-level agreement was assessed using the root mean square error (RMSE), Pearson correlation coefficient (Pearson's rho), and mean coherence. Acoustic features such as fundamental frequency (F0), jitter, shimmer, noise-to-harmonic ratio (NHR), cepstral peak prominence (CPP), and cepstral/spectral index of dysphonia (CSID) were extracted and compared for correlation analysis and intraclass correlation coefficients (ICC).
Signal-domain analysis revealed favorable agreement between FUDAN VOICE and CSL, with low RMSE and preserved spectral structure below 5 kHz, although high-frequency attenuation (> 5 kHz) was observed. Feature-level verification demonstrated excellent concordance for F0 mean, CPP, and CSID (r > 0.82, ICC > 0.82) across environments, whereas F0 extreme values-particularly F0 min-showed marked degradation in non-soundproof settings. Jitter and NHR remained robust, while shimmer exhibited environmental sensitivity.
FUDAN VOICE achieves reliable remote acoustic acquisition for F0 mean, CPP, and CSID across soundproof and office environments, though high-frequency attenuation (> 5 kHz) and noisy-setting F0 degradation warrant caution. This study establishes the first validation framework for super-app-based voice acquisition, supporting the standardization of WeChat mini-programs in mobile health.

PMID:
42760223
Bibliographic data and abstract were imported from PubMed on 19 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 15
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement