arXiv:2412.09296cs.CV2024-12中稿 · AAAI

让任意人脸视频开口说话,眼神自然、表情真实、动作有节奏。

GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic Expression

  • 用潜空间导航提升跨风格泛化能力,分离身份与动作,修正眼神异常。
  • 通过语音韵律感知的扩散模型生成符合语调的头部姿态。
  • 分阶段训练分离唇动与眨眼等动作,有限数据下仍生成逼真表情。

音频驱动的人脸动画需在多样输入人脸和音频-面部动作复杂关联下实现音画无缝融合。为此,我们提出稳健框架GoHD,可从任意参考身份生成高度真实、富有表现力且可控的肖像视频。GoHD创新性地引入三个核心模块:首先,采用潜空间导航的动画模块,增强对未见风格的泛化能力,实现运动与身份的高解耦,并引入注视方向以修正此前被忽视的不自然眼动;其次,设计基于Conformer结构的条件扩散模型,确保头姿对语音韵律的感知;第三,针对有限训练数据下从音频估计唇同步与真实表情的问题,提出两阶段训练策略,将频繁且逐帧的唇部运动蒸馏与更依赖时间但较少受音频影响的动作(如眨眼、皱眉)解耦。大量实验验证了GoHD卓越的泛化能力,在任意主体上均能生成高质量对话人脸结果。

原文摘要 · Abstract (English)

Audio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In response, we propose a robust framework GoHD designed to produce highly realistic, expressive, and controllable portrait videos from any reference identity with any motion. GoHD innovates with three key modules: Firstly, an animation module utilizing latent navigation is introduced to improve the generalization ability across unseen input styles. This module achieves high disentanglement of motion and identity, and it also incorporates gaze orientation to rectify unnatural eye movements that were previously overlooked. Secondly, a conformer-structured conditional diffusion model is designed to guarantee head poses that are aware of prosody. Thirdly, to estimate lip-synchronized and realistic expressions from the input audio within limited training data, a two-stage training strategy is devised to decouple frequent and frame-wise lip motion distillation from the generation of other more temporally dependent but less audio-related motions, e.g., blinks and frowns. Extensive experiments validate GoHD's advanced generalization capabilities, demonstrating its effectiveness in generating realistic talking face results on arbitrary subjects.

人脸动画音频驱动扩散模型解耦表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。