Skip to main navigation Skip to search Skip to main content

Sound event detection with audio-text models and heterogeneous temporal annotations

Research output: Chapter in Book/Report/Conference proceedingConference contributionScientificpeer-review

10 Downloads (Pure)

Abstract

Recent advances in generating synthetic captions based on audio and related metadata allow using the information contained in natural language as input for other audio tasks. In this paper, we propose a novel method to guide a sound event detection system with free-form text. We use machine-generated captions as complementary information to the strong labels for training, and evaluate the systems using different types of textual inputs. In addition, we study a scenario where only part of the training data has strong labels, and the rest of it only has temporally weak labels. Our findings show that synthetic captions improve the performance in both cases compared to the CRNN architecture typically used for sound event detection. On a dataset of 50 highly unbalanced classes, the PSDS-1 score increases from 0.223 to 0.277 when trained with strong labels, and from 0.166 to 0.218 when half of the training data has only weak labels.
Original languageEnglish
Title of host publication2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)
PublisherIEEE
ISBN (Electronic)979-8-3315-3745-6
DOIs
Publication statusPublished - 2025
Publication typeA4 Article in conference proceedings
EventIEEE Workshop on Applications of Signal Processing to Audio and Acoustics - Tahoe City, United States
Duration: 12 Oct 202515 Oct 2025

Publication series

NameIEEE Workshop on Applications of Signal Processing to Audio and Acoustics
ISSN (Electronic)1947-1629

Conference

ConferenceIEEE Workshop on Applications of Signal Processing to Audio and Acoustics
Country/TerritoryUnited States
City Tahoe City
Period12/10/2515/10/25

Funding

This work was supported by Academy of Finland grant 332063 "Teaching machines to listen". The authors wish to thank CSC-IT Centre of Science Ltd., Finland, for providing computational resources.

Publication forum classification

  • Publication forum level 1

Fingerprint

Dive into the research topics of 'Sound event detection with audio-text models and heterogeneous temporal annotations'. Together they form a unique fingerprint.

Cite this