arXiv:2412.00733cs.CVcs.GR2024-12CVPR被引 137

用视频扩散变压器实现高动态真实人像动画,解决视角变化与背景复杂问题。

Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

论文配图:Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
图 1 · 摘自论文原文
  • 采用基于Transformer的视频生成模型,结合因果3D VAE与堆叠Transformer层保持身份一致。
  • 在新构建的野外数据集上,生成视频在多角度、动态背景下的真实感显著优于旧方法。
  • 适合做高质量人像动画、影视特效或虚拟角色交互的研究者与开发者参考。

现有肖像图像动画方法在处理非正对视角、生成肖像周围的动态物体以及创造沉浸式真实背景方面仍面临挑战。本文首次应用预训练的基于Transformer的视频生成模型,展现出强泛化能力,可生成高度动态且逼真的肖像动画视频,有效解决上述难题。由于采用新型视频骨干网络,以往基于U-Net的身份保持、音频条件控制和视频外推方法不再适用。为此,我们设计了一个身份参考网络,由因果3D VAE与多层堆叠Transformer组成,确保视频序列中面部身份的一致性。同时,我们研究了多种语音音频条件机制与运动帧生成策略,实现由语音驱动的连续视频生成。实验在基准数据集和新提出的野外数据集上验证,结果表明本方法在多样化视角、动态沉浸场景下生成的肖像视频明显优于先前方法。可视化效果与源代码见:https://fudan-generative-vision.github.io/hallo3/

原文摘要 · Abstract (English)

Existing methodologies for animating portrait images face significant challenges, particularly in handling non-frontal perspectives, rendering dynamic objects around the portrait, and generating immersive, realistic backgrounds. In this paper, we introduce the first application of a pretrained transformer-based video generative model that demonstrates strong generalization capabilities and generates highly dynamic, realistic videos for portrait animation, effectively addressing these challenges. The adoption of a new video backbone model makes previous U-Net-based methods for identity maintenance, audio conditioning, and video extrapolation inapplicable. To address this limitation, we design an identity reference network consisting of a causal 3D VAE combined with a stacked series of transformer layers, ensuring consistent facial identity across video sequences. Additionally, we investigate various speech audio conditioning and motion frame mechanisms to enable the generation of continuous video driven by speech audio. Our method is validated through experiments on benchmark and newly proposed wild datasets, demonstrating substantial improvements over prior methods in generating realistic portraits characterized by diverse orientations within dynamic and immersive scenes. Further visualizations and the source code are available at: https://fudan-generative-vision.github.io/hallo3/.

人像动画视频生成扩散模型Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。