Abstract
Pre-trained speech representations like wav2vec 2.0 are a powerful tool for automatic speech recognition (ASR). Yet many endangered languages lack sufficient data for pre-training such models, or are predominantly oral vernaculars without a standardised writing system, precluding fine-tuning. Query-by-example spoken term detection (QbE-STD) offers an alternative for iteratively indexing untranscribed speech corpora by locating spoken query terms. Using data from 7 Australian Aboriginal languages and a regional variety of Dutch, all of which are endangered or vulnerable, we show that QbE-STD can be improved by leveraging representations developed for ASR (wav2vec 2.0: the English monolingual model and XLSR53 multilingual model). Surprisingly, the English model outperformed the multilingual model on 4 Australian language datasets, raising questions around how to optimally leverage self-supervised speech representations for QbE-STD. Nevertheless, we find that wav2vec 2.0 representations (either English or XLSR53) offer large improvements (56-86% relative) over state-of-the-art approaches on our endangered language datasets.
| Original language | English |
|---|---|
| Title of host publication | ASRU 2021: 2021 IEEE Automatic Speech Recognition and Understanding Workshop |
| Subtitle of host publication | proceedings |
| Place of Publication | Piscataway, NJ |
| Publisher | Institute of Electrical and Electronics Engineers (IEEE) |
| Pages | 1094-1101 |
| Number of pages | 8 |
| ISBN (Electronic) | 9781665437394 |
| DOIs | |
| Publication status | Published - 2021 |
| Externally published | Yes |
| Event | 2021 IEEE Automatic Speech Recognition and Understanding Workshop - Cartagena, Colombia Duration: 13 Dec 2021 → 17 Dec 2021 |
Conference
| Conference | 2021 IEEE Automatic Speech Recognition and Understanding Workshop |
|---|---|
| Abbreviated title | ASRU 2021 |
| Country/Territory | Colombia |
| City | Cartagena |
| Period | 13/12/21 → 17/12/21 |
Keywords
- endangered languages
- feature extraction
- language documentation
- spoken term detection
Fingerprint
Dive into the research topics of 'Leveraging pre-trained representations to improve access to untranscribed speech from endangered languages'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver