arXiv:2409.09326cs.CV2024-09

用局部仿射变形提升语音驱动唇动的逼真度与连贯性

LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping Deformation

  • 通过局部仿射变形场建模音频驱动的复杂唇部运动
  • 在LRS3数据集上实现8.7%的面部动作精度提升
  • 适合虚拟主播、AI客服等需要自然唇动的应用场景

在真实感虚拟形象生成领域,语音驱动唇部运动的保真度对实现自然交互至关重要。现有方法面临两大挑战:生成唇部姿态多样性不足导致表现力欠缺,以及因时间连贯性差引发的明显失真运动。为此,本文提出LawDNet,一种基于局部仿射变形机制的新型深度学习架构。该机制通过可控制的非线性变形场,建模音频输入下复杂的唇部运动,其核心为聚焦于深层特征图抽象关键点的局部仿射变换,提供了一种网络中特征变形的新通用范式。此外,LawDNet采用双流判别器以增强帧间连续性,并引入人脸归一化技术应对姿态和场景变化。大量实验表明,相比以往方法,LawDNet在鲁棒性和唇动动态表现上均具显著优势。本文提出的算法、训练数据、源代码及预训练模型将向研究社区开放。

原文摘要 · Abstract (English)

In the domain of photorealistic avatar generation, the fidelity of audio-driven lip motion synthesis is essential for realistic virtual interactions. Existing methods face two key challenges: a lack of vivacity due to limited diversity in generated lip poses and noticeable anamorphose motions caused by poor temporal coherence. To address these issues, we propose LawDNet, a novel deep-learning architecture enhancing lip synthesis through a Local Affine Warping Deformation mechanism. This mechanism models the intricate lip movements in response to the audio input by controllable non-linear warping fields. These fields consist of local affine transformations focused on abstract keypoints within deep feature maps, offering a novel universal paradigm for feature warping in networks. Additionally, LawDNet incorporates a dual-stream discriminator for improved frame-to-frame continuity and employs face normalization techniques to handle pose and scene variations. Extensive evaluations demonstrate LawDNet's superior robustness and lip movement dynamism performance compared to previous methods. The advancements presented in this paper, including the methodologies, training data, source codes, and pre-trained models, will be made accessible to the research community.

语音驱动唇动合成仿射变形虚拟形象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。