arXiv:2508.18734cs.CVcs.AI2025-08中稿 · IEEE ASRU 2025被引 4

通过动态调整音视频特征权重,提升噪声环境下语音识别准确率

Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion

  • 基于音频失真评分,按令牌级自适应加权音视频特征
  • 在LRS3数据集上相对误差率降低16.51%至42.67%
  • 适合追求高鲁棒性音视频识别的系统开发者

噪声环境下鲁棒的音视频语音识别(AVSR)仍具挑战,现有系统难以评估音频可靠性并动态调整模态依赖。本文提出路由门控跨模态特征融合框架,根据令牌级声学退化分数自适应重加权音视频特征。通过基于音视频特征融合的路由机制,模型在解码层中对不可靠音频令牌降权,并通过门控交叉注意力增强视觉线索,实现音频质量下降时向视觉模态的动态切换。在LRS3数据集上的实验表明,该方法相较AV-HuBERT实现16.51%至42.67%的相对词错误率降低。消融实验验证了路由与门控机制对真实噪声场景下鲁棒性的贡献。

原文摘要 · Abstract (English)

Robust audio-visual speech recognition (AVSR) in noisy environments remains challenging, as existing systems struggle to estimate audio reliability and dynamically adjust modality reliance. We propose router-gated cross-modal feature fusion, a novel AVSR framework that adaptively reweights audio and visual features based on token-level acoustic corruption scores. Using an audio-visual feature fusion-based router, our method down-weights unreliable audio tokens and reinforces visual cues through gated cross-attention in each decoder layer. This enables the model to pivot toward the visual modality when audio quality deteriorates. Experiments on LRS3 demonstrate that our approach achieves an 16.51-42.67% relative reduction in word error rate compared to AV-HuBERT. Ablation studies confirm that both the router and gating mechanism contribute to improved robustness under real-world acoustic noise.

音视频识别噪声鲁棒跨模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。