arXiv:2603.07263cs.SDeess.AS2026-03

让语音识别模型学会看视频上下文,提升复杂场景下的识别准确率。

Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning

  • 构建视听思维链,强制声画信号在中间阶段对齐
  • 在多个数据集上实现当前最优性能,有效缓解单一模态依赖
  • 开源数据与代码,适合多模态语音研究者使用

音频-视觉语音识别(AVSR)是将视觉信息融入语音识别的扩展方法。现有方法主要关注唇动,忽略了视频中丰富的上下文信息,如说话场景和屏幕文字。为解决此类包含丰富视觉上下文的AVSR(CAVSR),我们提出VASR,使其能够“看见”并推理视觉上下文以提升语音识别效果。具体而言,我们构建了视听思维链(AV-CoT),显式地在声学信号与视觉证据之间强制进行中间阶段的跨模态对齐。这种以证据驱动的推理机制有效缓解了‘单模态主导’问题,即模型过度依赖视觉或未能利用视觉信息。此外,为应对数据稀缺问题,我们构建并发布了相应的数据流水线与测试集。实验表明,AV-CoT能有效缓解单模态主导问题,在CAVSR任务中达到当前最优表现。项目已开源。

原文摘要 · Abstract (English)

Audio-visual speech recognition (AVSR) is an extension of ASR that incorporates visual signals. Current AVSR approaches primarily focus on lip motion, largely overlooking rich context present in the video such as speaking scene and on-screen text. To tackle such CAVSR (AVSR including rich visual Context), we propose VASR designed to "see" and reason the visual context to improve speech recognition. Specifically, we construct an Audio-Visual Chain-of-Thought (AV-CoT) that explicitly enforces intermediate cross-modal grounding between acoustic signals and visual evidence. This evidence-driven reasoning mitigates the "single-modality dominance" problem, where models either over-rely on visual context or fail to utilize it. Besides, to address the data scarcity, we construct and release a corresponding data pipeline and test set. Experiments show that AV-CoT effectively mitigates the single-modality dominance, achieving state-of-the-art performance in CAVSR. The project is open-sourced.

多模态语音识别视觉上下文思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。