通过数据增强与识别引导的CTC损失,提升多语种语音模型在少样本下的性能。
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
- 采用数据增强和识别感知的CTC损失优化微调策略。
- 在少样本设置下,语音识别错误率降低30%,语言识别准确率提升14%。
- 适合需要高效利用少量标注数据的多语种语音任务研究者。
基于自监督或有监督预训练的多语种语音基础模型(SFM)在语言识别(LID)和自动语音识别(ASR)等任务上表现强劲,但在微调阶段受限于资源时性能下降。本文针对ML-SUPERB 2.0中的多语种LID与ASR任务,探索了多种适应SFM的策略,包括冻结上游训练、部分微调及低秩适配。此外,引入数据增强以缓解少样本场景下的性能差距,并提出基于语言识别的连接时序分类(LID-Aware CTC)损失进行正则化。实验结果表明,该方法在ML-SUPERB 2.0上相较基线实现14%相对提升的LID准确率,以及30%相对降低的ASR词错误率(CER),并在Interspeech 2025 ML-SUPERB 2.0挑战赛中获得第二名。
原文摘要 · Abstract (English)
Multilingual speech processing with self-supervised or supervised pre-trained Speech Foundation Models (SFM) has achieved strong performance on tasks like Language Identification (LID) and Automatic Speech Recognition (ASR). However, these models struggle with limited resources during fine-tuning. This paper enhances multilingual LID and ASR on ML-SUPERB 2.0 by exploring multiple strategies for adapting SFMs, including frozen upstream training, partial fine-tuning, and low-rank adaptation. Furthermore, we employ data augmentation to mitigate performance gaps in few-shot settings and introduce LID Connectionist Temporal Classification (CTC) loss for regularization. Our approach achieves a 14% relative improvement in LID accuracy and a 30% relative reduction in ASR CER over the baseline on ML-SUPERB 2.0, securing second place in the Interspeech 2025 ML-SUPERB 2.0 Challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。