Siirry päänavigointiin Siirry hakuun Siirry pääsisältöön

Multi-Label Zero-Shot Audio Classification with Temporal Attention

Tutkimustuotos: KonferenssiartikkeliTieteellinenvertaisarvioitu

2 Sitaatiot (Scopus)
40 Lataukset (Pure)

Abstrakti

Zero-shot learning models are capable of classifying new classes by transferring knowledge from the seen classes using auxiliary information. While most of the existing zero-shot learning methods focused on single-label classification tasks, the present study introduces a method to perform multi-label zero-shot audio classification. To address the challenge of classifying multi-label sounds while generalizing to unseen classes, we adapt temporal attention. The temporal attention mechanism assigns importance weights to different audio segments based on their acoustic and semantic compatibility, thus enabling the model to capture the varying dominance of different sound classes within an audio sample by focusing on the segments most relevant for each class. This leads to more accurate multi-label zero-shot classification than methods employing temporally aggregated acoustic features without weighting, which treat all audio segments equally. We evaluate our approach on a subset of AudioSet against a zero-shot model using uniformly aggregated acoustic features, a zero-rule baseline, and the proposed method in the supervised scenario. Our results show that temporal attention enhances the zero-shot audio classification performance in multi-label scenario.

AlkuperäiskieliEnglanti
Otsikko2024 18th International Workshop on Acoustic Signal Enhancement, IWAENC 2024 - Proceedings
KustantajaIEEE
Sivut250-254
Sivumäärä5
ISBN (elektroninen)979-8-3503-6185-8
DOI - pysyväislinkit
TilaJulkaistu - 2024
OKM-julkaisutyyppiA4 Artikkeli konferenssijulkaisussa
TapahtumaInternational Workshop on Acoustic Signal Enhancement - Aalborg, Tanska
Kesto: 9 syysk. 202412 syysk. 2024

Julkaisusarja

Nimi
ISSN (painettu)2639-4316
ISSN (elektroninen)2835-3439

Conference

ConferenceInternational Workshop on Acoustic Signal Enhancement
Maa/AlueTanska
KaupunkiAalborg
Ajanjakso9/09/2412/09/24

Julkaisufoorumi-taso

  • Jufo-taso 1

!!ASJC Scopus subject areas

  • Signal Processing
  • Acoustics and Ultrasonics

Sormenjälki

Sukella tutkimusaiheisiin 'Multi-Label Zero-Shot Audio Classification with Temporal Attention'. Ne muodostavat yhdessä ainutlaatuisen sormenjäljen.

Siteeraa tätä