通过多粒度一致性提升语音识别在噪声下的鲁棒性
MGSC: A Multi-granularity Consistency Framework for Robust End-to-end Asr
- 引入多粒度软一致性框架,同时约束句子语义和词元对齐
- 在多种噪声条件下字符错误率相对降低8.7%
- 适合需要高可靠性语音识别的场景
端到端语音识别模型虽在基准测试中表现优异,但在噪声环境下常产生灾难性语义错误。我们将其脆弱性归因于主流的‘直接映射’目标函数——仅惩罚最终输出错误,而未约束模型内部计算过程。为此,我们提出多粒度软一致性(MGSC)框架,一种无需修改模型结构、可即插即用的模块,通过同时正则化宏观层面的句子语义与微观层面的词元对齐,强制模型内部自一致性。关键的是,我们的工作首次揭示了这两种一致性粒度间的强大协同效应:联合优化带来的鲁棒性提升显著超过各自独立贡献之和。在公开数据集上,MGSC在多种噪声条件下平均字符错误率相对降低8.7%,主要防止了严重改变语义的错误。研究证明,强制内部一致性是构建更鲁棒可信AI的关键一步。
原文摘要 · Abstract (English)
End-to-end ASR models, despite their success on benchmarks, often pro-duce catastrophic semantic errors in noisy environments. We attribute this fragility to the prevailing 'direct mapping' objective, which solely penalizes final output errors while leaving the model's internal computational pro-cess unconstrained. To address this, we introduce the Multi-Granularity Soft Consistency (MGSC) framework, a model-agnostic, plug-and-play module that enforces internal self-consistency by simultaneously regulariz-ing macro-level sentence semantics and micro-level token alignment. Cru-cially, our work is the first to uncover a powerful synergy between these two consistency granularities: their joint optimization yields robustness gains that significantly surpass the sum of their individual contributions. On a public dataset, MGSC reduces the average Character Error Rate by a relative 8.7% across diverse noise conditions, primarily by preventing se-vere meaning-altering mistakes. Our work demonstrates that enforcing in-ternal consistency is a crucial step towards building more robust and trust-worthy AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。