通过潜在空间混合提升语音多样性,改善低资源语言识别效果
Bridging the Language Gap: Synthetic Voice Diversity via Latent Mixup for Equitable Speech Recognition
- 在潜在空间混合法合成多样化语音数据
- 显著提升低资源语言语音识别准确率
- 适合关注语音技术公平性的研究者
当前语音任务的机器学习模型在英语等高资源语言上表现优异,主要得益于充足的数据。这种数据差异导致低资源语言面临性能不公,因数据收集困难且成本高昂。本文提出一种新型语音语料数据增强技术,旨在缓解这一差距。通过全面实验,证明该方法能显著提升低资源语言自动语音识别系统的性能。此外,相比现有增强策略,本方法表现更优,为提升代表性不足语言群体的语音技术提供了实用解决方案。
原文摘要 · Abstract (English)
Modern machine learning models for audio tasks often exhibit superior performance on English and other well-resourced languages, primarily due to the abundance of available training data. This disparity leads to an unfair performance gap for low-resource languages, where data collection is both challenging and costly. In this work, we introduce a novel data augmentation technique for speech corpora designed to mitigate this gap. Through comprehensive experiments, we demonstrate that our method significantly improves the performance of automatic speech recognition systems on low-resource languages. Furthermore, we show that our approach outperforms existing augmentation strategies, offering a practical solution for enhancing speech technology in underrepresented linguistic communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。