arXiv:2501.04204cs.CVcs.MM2025-01中稿 · presentation at IC…被引 13

用语音驱动生成逼真唇动视频,提升真实场景下唇读模型鲁棒性。

LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition

  • 基于发音单元(viseme)生成语音驱动的合成唇动视频。
  • 在LRW数据集上超越当前最优模型,尤其在复杂条件下优势更明显。
  • 引入视觉注意力机制,增强模型对关键语音时段的识别能力。

视觉语音识别(VSR),即唇读,因其广泛应用而受到广泛关注。深度学习与硬件进步显著提升了唇读模型性能。然而,现有数据集多为稳定录制,唇部动作变化有限,导致模型在真实场景中敏感度高。为此,我们提出LipGen框架,通过语音驱动的合成视觉数据提升模型鲁棒性,缓解数据局限。同时引入辅音分类辅助任务,结合注意力机制,有效整合时间信息,引导模型关注语音相关片段,增强判别能力。实验表明,该方法在唇读在野外(LRW)数据集上优于当前最先进水平,且在挑战性条件下表现更为突出。

原文摘要 · Abstract (English)

Visual speech recognition (VSR), commonly known as lip reading, has garnered significant attention due to its wide-ranging practical applications. The advent of deep learning techniques and advancements in hardware capabilities have significantly enhanced the performance of lip reading models. Despite these advancements, existing datasets predominantly feature stable video recordings with limited variability in lip movements. This limitation results in models that are highly sensitive to variations encountered in real-world scenarios. To address this issue, we propose a novel framework, LipGen, which aims to improve model robustness by leveraging speech-driven synthetic visual data, thereby mitigating the constraints of current datasets. Additionally, we introduce an auxiliary task that incorporates viseme classification alongside attention mechanisms. This approach facilitates the efficient integration of temporal information, directing the model's focus toward the relevant segments of speech, thereby enhancing discriminative capabilities. Our method demonstrates superior performance compared to the current state-of-the-art on the lip reading in the wild (LRW) dataset and exhibits even more pronounced advantages under challenging conditions.

唇读视频生成语音驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。