用噪声自适应编码与动态融合提升语音识别纠错在嘈杂环境下的效果
Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition
- 通过噪声自适应声学编码器增强模型对不同噪声的适应性
- 引入异构特征补偿动态融合机制,提升多模态信息利用效率
- 结合强化学习训练,显著提升未见噪声场景下的泛化能力
近年来,大语言模型(LLM)在自动语音识别(ASR)后处理中的生成错误纠正(GER)任务中取得显著进展。然而,在复杂噪声环境下,仍面临适应性差、信息利用率低等问题,导致GER效果受限。为此,本文提出一种抗噪声多模态GER框架(Denoising GER)。该框架通过噪声自适应声学编码器提升模型对各类噪声场景的适应能力,并利用异构特征补偿动态融合(HFCDF)机制优化多模态信息整合,增强LLM对多模态信息的利用。此外,引入强化学习(RL)训练策略以提升模型预测能力。实验结果表明,Denoising GER在噪声环境中显著提升了准确率与鲁棒性,并在未见过的噪声场景中展现出良好的泛化能力。
原文摘要 · Abstract (English)
In recent years, large language models (LLM) have made significant progress in the task of generation error correction (GER) for automatic speech recognition (ASR) post-processing. However, in complex noisy environments, they still face challenges such as poor adaptability and low information utilization, resulting in limited effectiveness of GER. To address these issues, this paper proposes a noise-robust multi-modal GER framework (Denoising GER). The framework enhances the model's adaptability to different noisy scenarios through a noise-adaptive acoustic encoder and optimizes the integration of multi-modal information via a heterogeneous feature compensation dynamic fusion (HFCDF) mechanism, improving the LLM's utilization of multi-modal information. Additionally, reinforcement learning (RL) training strategies are introduced to enhance the model's predictive capabilities. Experimental results demonstrate that Denoising GER significantly improves accuracy and robustness in noisy environments and exhibits good generalization abilities in unseen noise scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。