用中间CTC监督的专家混合模型提升口音语音识别准确率
Mixture-of-Experts with Intermediate CTC Supervision for Accented Speech Recognition
- 专家混合架构结合中间CTC监督,促进专家分工与泛化
- 在低/高资源条件下,对已见和未见口音均实现最高29.3%误识率降低
- 适合需要跨口音鲁棒性的语音识别系统开发者
口音语音识别仍是自动语音识别(ASR)的挑战,因多数模型基于少数高资源英语口音训练,导致其他口音性能显著下降。无口音依赖方法虽增强鲁棒性,但对重口音或未见口音表现不佳;而口音特定方法依赖有限且常含噪声的标签。我们提出Moe-Ctc:一种带中间CTC监督的专家混合架构,联合促进专家专精与泛化。训练时,口音感知路由引导专家捕捉口音特异性模式,推理时渐进过渡至无标签路由。每个专家配备独立的CTC头以对齐路由与转录质量,路由增强损失进一步稳定优化。在Mcv-Accent基准上,无论低资源或高资源条件,对已见及未见口音均实现持续提升,相比强基线FastConformer最高降低29.3%相对词错误率(WER)。
原文摘要 · Abstract (English)
Accented speech remains a persistent challenge for automatic speech recognition (ASR), as most models are trained on data dominated by a few high-resource English varieties, leading to substantial performance degradation for other accents. Accent-agnostic approaches improve robustness yet struggle with heavily accented or unseen varieties, while accent-specific methods rely on limited and often noisy labels. We introduce Moe-Ctc, a Mixture-of-Experts architecture with intermediate CTC supervision that jointly promotes expert specialization and generalization. During training, accent-aware routing encourages experts to capture accent-specific patterns, which gradually transitions to label-free routing for inference. Each expert is equipped with its own CTC head to align routing with transcription quality, and a routing-augmented loss further stabilizes optimization. Experiments on the Mcv-Accent benchmark demonstrate consistent gains across both seen and unseen accents in low- and high-resource conditions, achieving up to 29.3% relative WER reduction over strong FastConformer baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。