Skip to main navigation Skip to search Skip to main content

Syllable Discovery and Cross-Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model

  • Puyuan Peng
  • , Shang Wen Li
  • , Okko Räsänen
  • , Abdelrahman Mohamed
  • , David Harwath

Research output: Chapter in Book/Report/Conference proceedingConference contributionScientificpeer-review

14 Citations (Scopus)
8 Downloads (Pure)

Abstract

In this paper, we show that representations capturing syllabic units emerge when training a self-supervised speech model with a visually-grounded training objective. We demonstrate that a nearly identical model architecture (HuBERT) trained with a masked language modeling loss does not exhibit this same ability, suggesting that the visual grounding objective is responsible for the emergence of this phenomenon. We propose the use of a minimum cut algorithm to automatically predict syllable boundaries in speech, followed by a 2-stage clustering method to group identical syllables together. We show that our model not only outperforms a state-of-the-art syllabic segmentation method on the language it was trained on (English), but also generalizes in a zero-shot fashion to Estonian. Finally, we show that the same model is capable of zero-shot generalization for a word segmentation task on 4 other languages from the Zerospeech Challenge, in some cases beating the previous state-of-the-art.

Original languageEnglish
Title of host publicationProceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
PublisherInternational Speech Communication Association
Pages391-395
Number of pages5
DOIs
Publication statusPublished - 2023
Publication typeA4 Article in conference proceedings
EventAnnual Conference of the International Speech Communication Association, INTERSPEECH - Dublin, Ireland
Duration: 20 Aug 202324 Aug 2023

Publication series

NameInterspeech
PublisherInternational Speech Communication Association
ISSN (Electronic)2958-1796

Conference

ConferenceAnnual Conference of the International Speech Communication Association, INTERSPEECH
Country/TerritoryIreland
CityDublin
Period20/08/2324/08/23

Keywords

  • self-supervised speech processing
  • speech segmentation
  • visually-grounded speech

Publication forum classification

  • Publication forum level 1

ASJC Scopus subject areas

  • Language and Linguistics
  • Human-Computer Interaction
  • Signal Processing
  • Software
  • Modelling and Simulation

Fingerprint

Dive into the research topics of 'Syllable Discovery and Cross-Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model'. Together they form a unique fingerprint.

Cite this