TY - GEN
T1 - Syllable Discovery and Cross-Lingual Generalization in a Visually Grounded, Self-Supervised Speech Model
AU - Peng, Puyuan
AU - Li, Shang Wen
AU - Räsänen, Okko
AU - Mohamed, Abdelrahman
AU - Harwath, David
N1 - Publisher Copyright:
© 2023 International Speech Communication Association. All rights reserved.
PY - 2023
Y1 - 2023
N2 - In this paper, we show that representations capturing syllabic units emerge when training a self-supervised speech model with a visually-grounded training objective. We demonstrate that a nearly identical model architecture (HuBERT) trained with a masked language modeling loss does not exhibit this same ability, suggesting that the visual grounding objective is responsible for the emergence of this phenomenon. We propose the use of a minimum cut algorithm to automatically predict syllable boundaries in speech, followed by a 2-stage clustering method to group identical syllables together. We show that our model not only outperforms a state-of-the-art syllabic segmentation method on the language it was trained on (English), but also generalizes in a zero-shot fashion to Estonian. Finally, we show that the same model is capable of zero-shot generalization for a word segmentation task on 4 other languages from the Zerospeech Challenge, in some cases beating the previous state-of-the-art.
AB - In this paper, we show that representations capturing syllabic units emerge when training a self-supervised speech model with a visually-grounded training objective. We demonstrate that a nearly identical model architecture (HuBERT) trained with a masked language modeling loss does not exhibit this same ability, suggesting that the visual grounding objective is responsible for the emergence of this phenomenon. We propose the use of a minimum cut algorithm to automatically predict syllable boundaries in speech, followed by a 2-stage clustering method to group identical syllables together. We show that our model not only outperforms a state-of-the-art syllabic segmentation method on the language it was trained on (English), but also generalizes in a zero-shot fashion to Estonian. Finally, we show that the same model is capable of zero-shot generalization for a word segmentation task on 4 other languages from the Zerospeech Challenge, in some cases beating the previous state-of-the-art.
KW - self-supervised speech processing
KW - speech segmentation
KW - visually-grounded speech
U2 - 10.21437/Interspeech.2023-2044
DO - 10.21437/Interspeech.2023-2044
M3 - Conference contribution
AN - SCOPUS:85164399698
T3 - Interspeech
SP - 391
EP - 395
BT - Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
PB - International Speech Communication Association
T2 - Annual Conference of the International Speech Communication Association, INTERSPEECH
Y2 - 20 August 2023 through 24 August 2023
ER -