arXiv:2606.05763eess.AScs.SD2026-06

提出多视角自监督框架,提升复杂场景下音视频语音识别鲁棒性。

M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition

论文配图:M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition
图 1 · 摘自论文原文
  • 设计多视角编码器学习视角不变的视觉语音表征。
  • 通过模态质量与同步建模,实现细粒度跨模态融合,性能提升29.4%。
  • 构建真实场景数据集AISHELL8-RealScene,助力未来鲁棒多模态研究。

音视频语音识别(AVSR)通过利用视觉线索增强语音识别鲁棒性,但实际场景中视角变化、音频失真和视觉遮挡导致模态质量下降和音视频异步加剧。本文提出一种新型模态感知多视角自监督表示框架(M2S-AVSR)。首先,引入多视角表示学习编码器,学习视角不变的视觉语音表征;其次,设计模态感知模块,显式建模模态质量与跨模态同步性,实现细粒度模态感知融合,支持解码阶段精细注入视觉信息。此外,发布公开数据集AISHELL8-RealScene,涵盖多种真实场景下的多视角对话数据,并建立语音识别基准。在英语与中文基准测试中,该方法在挑战性条件下表现优异:在LRS3上,视角扰动与视觉退化设置下相对提升达29.4%;在MISP2021-AVSR测试集上达到新最优性能;在户外场景的AISHELL8-RealScene上亦取得最佳结果。所提方法与数据集为真实环境下鲁棒语音与多模态任务研究提供有力支持。

原文摘要 · Abstract (English)

Audio-Visual Speech Recognition (AVSR) enhances speech recognition robustness by leveraging visual cues, while real-world scenarios remain challenging due to viewpoint variation, audio distortion, and visual occlusion, which degrade modality quality and increase audio-visual asynchrony. In this paper, we propose a novel Modality-aware Multi-view Self-supervised representation framework for robust Audio-Visual Speech Recognition (M2S-AVSR). First, we introduce a multi-view representation learning encoder to learn view-invariant visual speech representations. Next, we employ a modality-aware module that explicitly models modality quality and cross-modal synchrony to perform fine-grained modality-aware fusion, enabling fine-grained visual information injection during decoding. In addition, we release AISHELL8-RealScene, a public multi-scenario, multi-view conversational audio-visual dataset recorded in real-world environments, and establish a speech recognition benchmark on it. Experiments on English and Mandarin benchmarks demonstrate the effectiveness of the proposed method under challenging conditions. On LRS3, M2S-AVSR achieves up to 29.4% relative improvement under viewpoint perturbation and visual degradation settings. Our method also achieves new state-of-the-art performance on the MISP2021-AVSR test set. On AISHELL8-RealScene, it achieves the best result in outdoor scenes. The proposed method and dataset provide useful support for future research on robust speech and multimodal tasks under realistic conditions.

音视频识别多模态自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。