arXiv:2602.08293eess.AS2026-02中稿 · ed被引 1

用可学习的瓶颈令牌实现音视频信息高效融合,提升噪声下的语音识别鲁棒性。

Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition

  • 引入可学习的紧凑令牌作为跨模态桥梁,动态调节音视频信息流动。
  • 在低资源下超越基线模型,噪声适应融合使性能接近大规模系统。
  • 适合需要轻量级高鲁棒性的音视频语音识别场景,如智能设备应用。

音视频语音识别(AVSR)通过结合声学与视觉线索,在噪声环境下提升语音识别性能。核心挑战在于设计一种融合机制,使模型在音频信号受损时仍能有效利用视觉信息,同时保持对清晰语音的良好表现。本文提出CoBRA(Cross-modal Bottleneck for Robust AVSR),一种基于瓶颈的融合框架,引入一组紧凑的可学习令牌来中介跨模态交互。通过这些令牌调控信息流,即使在恶劣或域外噪声条件下,音频流也能可靠获取关键视觉线索。尽管训练数据有限,该模型仍超越同类基线,并通过噪声自适应融合保持与大规模系统相当的竞争力,展现出高效与鲁棒性。消融实验表明,融合深度是决定系统鲁棒性的最关键因素,凸显其在设计鲁棒性AVSR系统中的重要性。

原文摘要 · Abstract (English)

Audio-Visual Speech Recognition (AVSR) leverages both acoustic and visual cues to improve speech recognition under noisy conditions. A central question is how to design a fusion mechanism that allows the model to effectively exploit visual information when the audio signal is degraded, while maintaining strong performance on clean speech. We propose CoBRA (Cross-modal Bottleneck for Robust AVSR), a bottleneck-based fusion framework that introduces a compact set of learnable tokens to mediate cross-modal exchange. By regulating information flow through these tokens, the audio stream can reliably access essential visual cues even under adverse or out-of-domain noise. Despite limited training data, our model surpasses comparable baselines and remains competitive with large-scale systems through noise-adaptive fusion, demonstrating both efficiency and robustness. Ablation studies highlight that the depth of fusion is the most critical factor, underscoring its importance in designing robust AVSR systems.

音视频识别噪声鲁棒跨模态融合轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。