arXiv:2601.12436eess.AScs.AI2026-01中稿 · ICASSP2026被引 1

不用噪声掩码,用视频辅助净化音频,提升嘈杂环境下的语音识别准确率。

Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition

  • 用视觉信息隐式净化噪声音频,避免显式生成噪声掩码。
  • 在LRS3数据集上,噪声环境下识别错误率低于现有先进方法。
  • 适合需要高鲁棒性语音识别的场景,如智能客服、车载系统。

音视频语音识别(AVSR)通常通过融合抗噪的视觉线索与音频信号来提升嘈杂环境下的识别精度。然而,高噪声音频输入容易在特征融合过程中引入干扰。现有方法多采用基于掩码的策略,在特征交互时过滤噪声,但可能误删语义相关信息。本文提出一种端到端的噪声鲁棒型AVSR框架,结合语音增强技术,无需显式生成噪声掩码。该框架利用基于Conformer的瓶颈融合模块,借助视频信息隐式优化噪声音频特征。通过降低模态冗余并增强跨模态交互,有效保留语音语义完整性,实现稳健识别性能。在公开数据集LRS3上的实验表明,该方法在噪声条件下优于现有的先进掩码基基准模型。

原文摘要 · Abstract (English)

Audio-visual speech recognition (AVSR) typically improves recognition accuracy in noisy environments by integrating noise-immune visual cues with audio signals. Nevertheless, high-noise audio inputs are prone to introducing adverse interference into the feature fusion process. To mitigate this, recent AVSR methods often adopt mask-based strategies to filter audio noise during feature interaction and fusion, yet such methods risk discarding semantically relevant information alongside noise. In this work, we propose an end-to-end noise-robust AVSR framework coupled with speech enhancement, eliminating the need for explicit noise mask generation. This framework leverages a Conformer-based bottleneck fusion module to implicitly refine noisy audio features with video assistance. By reducing modality redundancy and enhancing inter-modal interactions, our method preserves speech semantic integrity to achieve robust recognition performance. Experimental evaluations on the public LRS3 benchmark suggest that our method outperforms prior advanced mask-based baselines under noisy conditions.

音视频识别语音增强鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。