arXiv:2510.22172cs.SDcs.CL2025-10

提升非自回归语音识别对齐精度,尤其改善英语法语表现

M-CIF: Multi-Scale Alignment For CIF-Based Non-Autoregressive ASR

  • 通过字符与音素级监督逐步融合到子词表示中实现多尺度对齐
  • 在CommonVoice数据集上,德语WER降4.21%,法语降3.05%
  • 音素和字符层级对增强对齐稳定性至关重要,适合多语言语音识别研究

连续积分-放电(CIF)机制为非自回归语音识别提供了有效的对齐方式,能生成平滑且单调的声学特征到目标标记映射,在中文任务上性能可媲美其他NAR方法。然而,缺乏细粒度引导时,其在英语、法语等语言上的稳定性下降。本文提出多尺度CIF(M-CIF),通过逐级将字符与音素级监督融入子词表示,实现多层级对齐,显著提升声学-文本对齐鲁棒性。实验表明,相比Paraformer基线,M-CIF在CommonVoice上使德语WER降低4.21%,法语降低3.05%。进一步引入音素混淆错误(PE)与空间分割错误(SE)作为评估指标,分析显示音素层与字符层对渐进式对齐具有关键作用。

原文摘要 · Abstract (English)

The Continuous Integrate-and-Fire (CIF) mechanism provides effective alignment for non-autoregressive (NAR) speech recognition. This mechanism creates a smooth and monotonic mapping from acoustic features to target tokens, achieving performance on Mandarin competitive with other NAR approaches. However, without finer-grained guidance, its stability degrades in some languages such as English and French. In this paper, we propose Multi-scale CIF (M-CIF), which performs multi-level alignment by integrating character and phoneme level supervision progressively distilled into subword representations, thereby enhancing robust acoustic-text alignment. Experiments show that M-CIF reduces WER compared to the Paraformer baseline, especially on CommonVoice by 4.21% in German and 3.05% in French. To further investigate these gains, we define phonetic confusion errors (PE) and space-related segmentation errors (SE) as evaluation metrics. Analysis of these metrics across different M-CIF settings reveals that the phoneme and character layers are essential for enhancing progressive CIF alignment.

语音识别非自回归多尺度对齐CIF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。