用声音特征相似性精选外语语音数据,提升低资源语音识别效果
Cross-lingual Data Selection Using Clip-level Acoustic Similarity for Enhancing Low-resource Automatic Speech Recognition
- 基于声学特征分布相似性,细粒度筛选外语语音片段
- 在低资源场景下显著提升识别准确率,可利用以往无效的外语数据
- 适合做多语言语音识别、尤其是资源匮乏语言的研究者
本文提出一种新型捐赠数据选择方法,用于提升低资源自动语音识别(ASR)性能。尽管高资源语言中ASR表现良好,但在数据有限的低资源场景下性能下降。常见做法是利用多语言自监督学习(SSL)模型和捐赠语言数据,但现有方法依赖语言级相似性,忽略语音片段级差异。为此,我们提出剪辑级声学标记分布相似性(CATDS),通过细粒度匹配目标语言与捐赠语言的声学特征分布,选出更相关的声音片段。该方法与SSL模型表示对齐,能选出更具挑战性但更有价值的样本。实验表明,CATDS优于传统选择方法,甚至可有效利用以往被认为有害的捐赠语言数据。
原文摘要 · Abstract (English)
This paper presents a novel donor data selection method to enhance low-resource automatic speech recognition (ASR). While ASR performs well in high-resource languages, its accuracy declines in low-resource settings due to limited training data. A common solution is to leverage multilingual self-supervised learning (SSL) models with donor languages. However, existing methods rely on language-level similarity, overlooking clip-level variations. To address this limitation, we propose clip-wise acoustic token distribution similarity (CATDS), a fine-grained selection method that identifies acoustically relevant donor clips for better alignment with the target language. Unlike existing clip-level selection methods, our method aligns with the representation of SSL models and offers more challenging yet valuable samples. Experimental results show that CATDS outperforms traditional selection methods and can even utilize donor languages previously considered detrimental.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。