Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Structured Prototype Learning with Feature Fusion for Sparse and Asynchronous Audio-Visual Depression Recognition.

Created on 15 Sep 2026

Authors

Zhonghui Jin, Pei He, Yangming Guo, Xiaodong Wang, Aiqing Fang

Published in

Sensors (Basel, Switzerland). Volume 26. Issue 17. Sep 02, 2026. Epub Sep 02, 2026.

Abstract

Audio-visual depression recognition in real-world scenarios is often challenged by temporal sparsity and cross-modal asynchrony, where depression-related cues may appear only in short segments and may not align precisely across modalities. Under such conditions, global pooling or dense attention tends to dilute sparse discriminative evidence with redundant context, leading to unstable utterance-level representations. To address this issue, we propose an audio-visual depression recognition framework that integrates modality feature adaptation, bidirectional cross-modal interaction, and graph-based prototype abstraction. Specifically, heterogeneous audio and visual streams are first transformed into compatible representations, after which bidirectional cross-attention models content-dependent dependencies across modalities without requiring index-wise correspondence. The fused tokens are then interpreted as graph nodes and aggregated into a compact set of semantic prototypes through graph convolution and differentiable soft clustering. In addition, audio perturbation is introduced during training as a task-oriented regularisation strategy for partial acoustic evidence loss and temporal misalignment. Experiments on the LMVD dataset demonstrate clear improvements on the primary depression-recognition task, while auxiliary evaluations on MIntRec and CMU-MOSI suggest that the structured prototype representation is beneficial for other temporally sparse audio-visual recognition tasks. These results indicate that structured prototype learning is effective for preserving sparse depression-related cues, while training-time perturbation provides a complementary regularisation effect under asynchronous multimodal conditions.

PMID:
42740199
Bibliographic data and abstract were imported from PubMed on 15 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 10
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement