arXiv:2410.14910eess.AScs.SD2024-10被引 3

用混合语音提升低资源语音识别,解决数据源不匹配问题

AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup

  • 通过混元对比增强,融合源域与目标域语音生成新样本
  • 仅需1小时单卡训练,11小时适配数据即超越基线
  • 适合语音识别数据少、跨场景迁移的实用场景

自监督学习(SSL)利用大量无标签数据学习丰富的语音表征,推动了自动语音识别(ASR)的发展,即使在仅有少量标注数据用于微调时也表现良好。然而,当预训练数据(源域)与微调数据(目标域)存在分布差异时,性能仍受挑战。为此,本文提出一种针对低资源ASR的新型域适应方法——AC-Mix(无偏对比混元),专注于联合嵌入架构。该方法通过在源域与目标域样本间插值生成混合数据视图,对SSL模型进行额外预训练以实现适配。实验表明,该方法在仅使用约11小时适配数据、单张GPU上仅需1小时训练时间的情况下,显著优于基线系统,且效果稳定。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) leverages large amounts of unlabelled data to learn rich speech representations, fostering improvements in automatic speech recognition (ASR), even when only a small amount of labelled data is available for fine-tuning. Despite the advances in SSL, a significant challenge remains when the data used for pre-training (source domain) mismatches the fine-tuning data (target domain). To tackle this domain mismatch challenge, we propose a new domain adaptation method for low-resource ASR focused on contrastive mixup for joint-embedding architectures named AC-Mix (agnostic contrastive mixup). In this approach, the SSL model is adapted through additional pre-training using mixed data views created by interpolating samples from the source and the target domains. Our proposed adaptation method consistently outperforms the baseline system, using approximately 11 hours of adaptation data and requiring only 1 hour of adaptation time on a single GPU with WavLM-Large.

语音识别自监督学习域适应低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。