Siirry päänavigointiin Siirry hakuun Siirry pääsisältöön

Corpus-based dialectometry with topic models

Tutkimustuotos: ArtikkeliTieteellinenvertaisarvioitu

23 Lataukset (Pure)

Abstrakti

This paper presents a topic modeling approach to corpus-based dialectometry. Topic models are most often used in text mining to find latent structure in a collection of documents. They are based on the idea that frequently co-occurring words present the same underlying topic. In this study, topic models are used on interview transcriptions containing dialectal speech directly, without any annotations or preselected features. The transcriptions are modeled on complete words, on character n-grams, and after automatical segmentation. Data from three languages, Finnish, Norwegian, and Swiss German, are scrutinized. The proposed method is capable of discovering clear dialectal differences in all three datasets, while reflecting the differences between them. The method provides a significant simplification of the dialectometric workflow, simultaneously saving time and increasing objectivity. Using the method on non-normalized data could also benefit text mining, which is the traditional field of topic modeling.
AlkuperäiskieliEnglanti
Sivut1-12
Sivumäärä12
JulkaisuJournal of Linguistic Geography
Vuosikerta12
Numero1
Varhainen verkossa julkaisun päivämäärä20 toukok. 2024
DOI - pysyväislinkit
TilaJulkaistu - 2024
OKM-julkaisutyyppiA1 Alkuperäisartikkeli tieteellisessä aikakauslehdessä

Julkaisufoorumi-taso

  • Jufo-taso 1

Sormenjälki

Sukella tutkimusaiheisiin 'Corpus-based dialectometry with topic models'. Ne muodostavat yhdessä ainutlaatuisen sormenjäljen.

Siteeraa tätä