用音视频对齐提升唇语识别准确率
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
- 通过音视频跨模态注意力捕捉全局对应关系
- 引入帧级局部对齐损失优化音频信息利用
- 在LRS2和CNVSRC.Single数据集上表现更优
唇语识别(VSR)旨在通过分析唇部动作的视觉信息来识别对应文本。由于唇部动作变化大且信息弱,VSR需有效利用各来源、各层次的信息。本文提出基于音视频跨模态对齐的VSR方法AlignVSR,将音频作为辅助信息源,利用音视频的全局与局部对应关系提升视觉到文本的推理能力。首先,通过视频帧到音频单元库的跨模态注意力机制捕获音视频全局对齐;其次,基于音视频时间对应性,引入帧级局部对齐损失以精化全局对齐,增强音频信息利用效率。在LRS2和CNVSRC.Single数据集上的实验结果一致表明,AlignVSR优于多个主流VSR方法,展现了卓越且稳健的性能。
原文摘要 · Abstract (English)
Visual Speech Recognition (VSR) aims to recognize corresponding text by analyzing visual information from lip movements. Due to the high variability and weak information of lip movements, VSR tasks require effectively utilizing any information from any source and at any level. In this paper, we propose a VSR method based on audio-visual cross-modal alignment, named AlignVSR. The method leverages the audio modality as an auxiliary information source and utilizes the global and local correspondence between the audio and visual modalities to improve visual-to-text inference. Specifically, the method first captures global alignment between video and audio through a cross-modal attention mechanism from video frames to a bank of audio units. Then, based on the temporal correspondence between audio and video, a frame-level local alignment loss is introduced to refine the global alignment, improving the utility of the audio information. Experimental results on the LRS2 and CNVSRC.Single datasets consistently show that AlignVSR outperforms several mainstream VSR methods, demonstrating its superior and robust performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。