arXiv:2509.16031cs.CV2025-09被引 1

提出GLip框架,提升唇语识别在光照、遮挡等复杂条件下的鲁棒性。

GLip: A Global-Local Integrated Progressive Framework for Robust Visual Speech Recognition

  • 分两阶段学习:先粗对齐全局与局部视觉特征,再融合上下文精修映射。
  • 在LRS2和LRS3上超越现有方法,新汉语数据集验证有效。
  • 擅长利用非遮挡局部区域的判别性线索,适合真实场景应用。

唇语识别(VSR)是从无声视频中识别语音的任务。尽管近年来取得显著进展,但多数方法对光照变化、遮挡、模糊和姿态改变等现实挑战关注不足。为此,我们提出GLip——一种全局-局部集成渐进式框架,以增强鲁棒性。该框架基于两点核心洞察:(i) 在不同条件下对视觉特征与语音内容进行初步粗对齐,有助于在复杂环境中学习精确的视觉-语音映射;(ii) 在恶劣条件下,某些局部区域(如未遮挡部分)往往比全局特征更具判别性。为此,GLip采用双路径特征提取结构,结合全局与局部特征,在两阶段渐进学习框架中实现。第一阶段利用易获取的音视频数据,学习全局与局部特征与对应语音单元的对齐,建立粗粒度但语义稳健的基础。第二阶段引入上下文增强模块(CEM),在时空维度上动态融合局部特征与相关全局上下文,将粗略表示精炼为精准映射。该框架通过渐进策略充分利用判别性局部区域,在多种视觉挑战下表现更优,并在LRS2和LRS3基准上持续领先。进一步在新提出的挑战性普通话数据集上验证了有效性。

原文摘要 · Abstract (English)

Visual speech recognition (VSR), also known as lip reading, is the task of recognizing speech from silent video. Despite significant advancements in VSR over recent decades, most existing methods pay limited attention to real-world visual challenges such as illumination variations, occlusions, blurring, and pose changes. To address these challenges, we propose GLip, a Global-Local Integrated Progressive framework designed for robust VSR. GLip is built upon two key insights: (i) learning an initial coarse alignment between visual features across varying conditions and corresponding speech content facilitates the subsequent learning of precise visual-to-speech mappings in challenging environments; (ii) under adverse conditions, certain local regions (e.g., non-occluded areas) often exhibit more discriminative cues for lip reading than global features. To this end, GLip introduces a dual-path feature extraction architecture that integrates both global and local features within a two-stage progressive learning framework. In the first stage, the model learns to align both global and local visual features with corresponding acoustic speech units using easily accessible audio-visual data, establishing a coarse yet semantically robust foundation. In the second stage, we introduce a Contextual Enhancement Module (CEM) to dynamically integrate local features with relevant global context across both spatial and temporal dimensions, refining the coarse representations into precise visual-speech mappings. Our framework uniquely exploits discriminative local regions through a progressive learning strategy, demonstrating enhanced robustness against various visual challenges and consistently outperforming existing methods on the LRS2 and LRS3 benchmarks. We further validate its effectiveness on a newly introduced challenging Mandarin dataset.

唇语识别多模态鲁棒性视觉语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。