Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

A Vision-Language Model as a Teacher for Bird Vocalization Detection

Created on 26 Sep 2026

Authors

Vengrovski, G., Gardner, T. J.

Abstract

Time-frequency detection of bird vocalizations is an important step toward turning weakly annotated field recordings into usable data for studying avian communication. Supervised detectors trained on human expert annotations scale poorly across species and recording conditions, so we introduce a teacher-student setup in which the teacher, a vision-language model, labels bounding boxes on spectrograms of citizen-scientist recordings, and those labels train a student, a self-supervised bioacoustic encoder. We find that the student generally exceeds both the teacher and supervised models trained on human annotations and performs strongly on held-out datasets for both time-frequency and onset-offset localization. YOLO detectors trained on our teacher labels match those trained on human annotations, suggesting that VLM labels can substitute for costly expert labeling.

Preprint server: bioRxiv
The authors list and abstract were imported from bioRxiv on 26 Sep 2026.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this preprint? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 14
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement