用稀疏对齐与视觉单元引导,提升语音识别在噪声下的鲁棒性。
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
- 通过稀疏对齐减少跨模态干扰,只关注关键视觉信息。
- 在0dB信噪比下相比基线提升37%相对性能。
- 适合做高噪声环境下的多模态语音识别研究者参考。
音频-视觉语音识别(AVSR)融合声学与视觉信息,以增强恶劣声学条件下的鲁棒性。近期大型语言模型(LLM)在自动语音识别中表现出色,并已应用于AVSR。然而,现有方法通常独立投影音视频特征或采用浅层融合,限制了跨模态对齐与互补交互,同时增加了LLM的计算负担。为此,我们提出AVUR-LLM,一种基于稀疏模态对齐与视觉单元引导精炼的LLM音频-视觉语音识别框架。在LRS3数据集上的实验表明,该方法达到当前AVSR最佳性能。在添加噪声条件下,0 dB信噪比时,相比基线系统实现37%的相对性能提升。
原文摘要 · Abstract (English)
Audio-Visual Speech Recognition (AVSR) integrates acoustic and visual information to enhance robustness in adverse acoustic conditions. Recent advances in Large Language Models (LLMs) have yielded competitive automatic speech recognition performance and shown effectiveness for AVSR. However, prior approaches project audio and visual features independently or apply shallow fusion, limiting cross-modal alignment and complementary exchange while increasing the LLM's computational load. To address this, we propose AVUR-LLM, an LLM-based Audio-Visual Speech Recognition via Sparse Modality Alignment and Visual Unit-Guided Refinement. Experiments on LRS3 demonstrate state-of-the-art results for AVSR. Under additive-noise conditions at 0 dB SNR, it achieves 37% relative improvement over the baseline system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。