提出双向增强框架,提升嘈杂环境下的音视频语音识别能力
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
- 采用不对称双流音频编码,构建跨模态交互基础
- 在LRS2/LRS3上超越现有方法,噪声下识别准确率显著提升
- 适合音视频融合、降噪场景的研究与应用
音视频语音识别(AVSR)通过融合视听模态提升识别效果,尤其在噪声环境下表现更优。然而,现有方法多采用单向增强或对称融合策略,难以捕捉视听数据的异质性与互补性,尤其在信息不对称条件下表现受限。为此,本文提出新型AVSR框架AD-AVSR,基于双向模态增强机制。首先引入音频双流编码策略,从多视角丰富音频表示,并有意构建不对称性以支持后续跨模态交互。增强过程包含两个核心组件:音频感知的视觉精炼模块(在音频引导下优化视觉表示),以及跨模态噪声抑制掩码模块(利用视觉线索优化音频表示),协同实现闭环双向信息流动。为进一步增强相关性鲁棒性,采用阈值选择机制过滤无关或弱相关视听配对。在LRS2和LRS3数据集上的大量实验表明,所提AD-AVSR在性能与噪声鲁棒性上持续优于当前最优方法,验证了模型设计的有效性。
原文摘要 · Abstract (English)
Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric fusion manner, which limits their capability to capture heterogeneous and complementary correlations of audio-visual data-especially under asymmetric information conditions. To tackle these gaps, we introduce a new AVSR framework termed AD-AVSR based on bidirectional modality enhancement. Specifically, we first introduce the audio dual-stream encoding strategy to enrich audio representations from multiple perspectives and intentionally establish asymmetry to support subsequent cross-modal interactions. The enhancement process involves two key components, Audio-aware Visual Refinement Module for enhanced visual representations under audio guidance, and Cross-modal Noise Suppression Masking Module which refines audio representations using visual cues, collaboratively leading to the closed-loop and bidirectional information flow. To further enhance correlation robustness, we adopt a threshold-based selection mechanism to filter out irrelevant or weakly correlated audio-visual pairs. Extensive experimental results on the LRS2 and LRS3 datasets indicate that our AD-AVSR consistently surpasses SOTA methods in both performance and noise robustness, highlighting the effectiveness of our model design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。