用真实语音训练桥梁模块,提升语音增强与识别的匹配度。
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
- 用真实噪声语音训练桥梁模块,避免模拟数据偏差。
- 通过多任务学习和WER反馈优化混合系数,降低识别错误率。
- 特别适合在真实场景下部署的语音识别系统优化。
单通道语音增强导致的信息损失会损害自动语音识别(ASR)性能。观测添加(OA)是一种有效的后处理方法,通过平衡噪声与增强语音来提升ASR效果,其关键在于确定合适的OA系数。然而,当前基于监督学习的桥梁模块仅使用模拟噪声语音训练,与真实噪声存在严重不匹配。本文提出基于真实噪声语音的训练策略:采用DNSMOS评估真实噪声语音的感知质量,无需对应干净语音标签;引入额外约束提升模块鲁棒性;通过ASR后端对不同OA系数下的每段语音计算词错误率(WER),构建多维向量,并结合多任务学习用于桥梁模块,以确定最优OA系数。在CHiME-4数据集上的实验表明,所提方法相比模拟数据训练的桥梁模块均有显著提升,尤其在真实测试集上表现更优。
原文摘要 · Abstract (English)
The information loss or distortion caused by single-channel speech enhancement (SE) harms the performance of automatic speech recognition (ASR). Observation addition (OA) is an effective post-processing method to improve ASR performance by balancing noisy and enhanced speech. Determining the OA coefficient is crucial. However, the currently supervised OA coefficient module, called the bridging module, only utilizes simulated noisy speech for training, which has a severe mismatch with real noisy speech. In this paper, we propose training strategies to train the bridging module with real noisy speech. First, DNSMOS is selected to evaluate the perceptual quality of real noisy speech with no need for the corresponding clean label to train the bridging module. Additional constraints during training are introduced to enhance the robustness of the bridging module further. Each utterance is evaluated by the ASR back-end using various OA coefficients to obtain the word error rates (WERs). The WERs are used to construct a multidimensional vector. This vector is introduced into the bridging module with multi-task learning and is used to determine the optimal OA coefficients. The experimental results on the CHiME-4 dataset show that the proposed methods all had significant improvement compared with the simulated data trained bridging module, especially under real evaluation sets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。