不训练模型,用语音可懂度自动调节降噪与识别的融合比例。
Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
- 通过后端识别模型直接估计可懂度,决定噪声语音与增强语音的融合权重。
- 在多个降噪-识别组合和数据集上均显著优于现有方法。
- 无需训练,通用性强,适合实际部署场景。
自动语音识别(ASR)在噪声环境下性能严重下降。尽管语音增强(SE)前端能有效抑制背景噪声,但常引入影响识别的失真。观测值添加(OA)通过融合噪声语音与增强语音,提升识别效果且无需修改SE或ASR模型参数。本文提出一种可懂度引导的OA方法,融合权重由后端ASR直接估算的可懂度得到。相比依赖训练神经预测器的已有方法,该方法无需训练,降低复杂度并增强泛化能力。在多种SE-ASR组合及数据集上的广泛实验表明,该方法具备强鲁棒性,并显著优于现有OA基线。对基于可懂度切换的替代方案及帧级与语句级OA的进一步分析也验证了设计的有效性。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) degrades severely in noisy environments. Although speech enhancement (SE) front-ends effectively suppress background noise, they often introduce artifacts that harm recognition. Observation addition (OA) addressed this issue by fusing noisy and SE enhanced speech, improving recognition without modifying the parameters of the SE or ASR models. This paper proposes an intelligibility-guided OA method, where fusion weights are derived from intelligibility estimates obtained directly from the backend ASR. Unlike prior OA methods based on trained neural predictors, the proposed method is training-free, reducing complexity and enhances generalization. Extensive experiments across diverse SE-ASR combinations and datasets demonstrate strong robustness and improvements over existing OA baselines. Additional analyses of intelligibility-guided switching-based alternatives and frame versus utterance-level OA further validate the proposed design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。