arXiv:2409.08872cs.CLcs.SD2024-09中稿 · O-COCOSDA 2025被引 2

用跨语言方法增强濒危语言语音识别数据,提升识别准确率

Exploring the Impact of Data Quantity on ASR in Extremely Low-resource Languages

  • 通过多语言语料库筛选与目标语言音系相近的语音片段
  • 在阿美语和赛德克语上实现显著的识别性能提升
  • 适合关注濒危语言数字化与低资源语音识别的研究者

本研究探讨了数据增强技术在极低资源自动语音识别(ASR)中的有效性,聚焦两种濒危南岛语——阿美语和赛德克语。鉴于自监督学习(SSL)在低资源场景下的潜力,我们考察了数据量对SSL模型持续预训练的影响。提出一种新颖的数据选择方案,利用多语言语料库,通过语言分类器提取语音嵌入,并采用一类分类器识别在语音和音系上与目标语言相近的语句。根据决策得分对语句排序并选取,确保在SSL-ASR流程中纳入高度相关数据。实验结果表明,该方法在阿美语和赛德克语上均取得显著性能提升,验证了通过跨语言迁移学习进行数据增强的可行性与前景。

原文摘要 · Abstract (English)

This study investigates the efficacy of data augmentation techniques for low-resource automatic speech recognition (ASR), focusing on two endangered Austronesian languages, Amis and Seediq. Recognizing the potential of self-supervised learning (SSL) in low-resource settings, we explore the impact of data volume on the continued pre-training of SSL models. We propose a novel data-selection scheme leveraging a multilingual corpus to augment the limited target language data. This scheme utilizes a language classifier to extract utterance embeddings and employs one-class classifiers to identify utterances phonetically and phonologically proximate to the target languages. Utterances are ranked and selected based on their decision scores, ensuring the inclusion of highly relevant data in the SSL-ASR pipeline. Our experimental results demonstrate the effectiveness of this approach, yielding substantial improvements in ASR performance for both Amis and Seediq. These findings underscore the feasibility and promise of data augmentation through cross-lingual transfer learning for low-resource language ASR.

语音识别低资源濒危语言数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。