用对比学习提升语音识别对口音的鲁棒性
Contrastive Regularization for Accent-Robust ASR

- 在CTC微调中加入句级对比损失,不改结构也不需口音标注
- 在未见口音下相对错误率降低25%~29%,多模型通用
- 让编码器表示更紧凑稳定,适合做口音鲁棒的语音系统
基于自监督声学预训练和CTC微调的语音识别系统在母语语音上表现优异,但对口音变化仍敏感。本文研究将监督对比学习(SupCon)作为轻量级、与口音无关的辅助目标用于CTC微调。通过句级对比损失正则化编码器表示,无需修改架构或显式口音监督。在L2-ARCTIC基准上的实验表明,多种预训练编码器均实现一致的词错误率(WER)下降,未见口音评估下相对减少达25%~29%。使用同句内余弦分散分析显示,SupCon在口音变化下促进更紧凑、更稳定的表示几何结构。总体而言,SupCon为提升口音鲁棒性提供了有效且模型无关的正则化策略。
原文摘要 · Abstract (English)
ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability. We investigate supervised contrastive learning (SupCon) as a lightweight, accent-invariant auxiliary objective for CTC fine-tuning. An utterance-level contrastive loss regularizes encoder representations without architectural modification or explicit accent supervision. Experiments on the L2-ARCTIC benchmark show consistent WER reductions across multiple pretrained encoders, with up to 25 -- 29\% relative reduction under unseen-accent evaluation. Analysis using within-transcript cosine dispersion indicates that SupCon promotes more compact and stable representation geometry under accent variability. Overall, SupCon provides an effective and model-agnostic regularization strategy for improving accent robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。