arXiv:2409.18083cs.CV2024-09ECCV被引 6

用3D人脸模型控制2D生成,实现逼真说话人脸视频

Stable Video Portraits

  • 结合2D扩散模型与3DMM,通过3D参数控制生成视频
  • 生成视频时序稳定,可零样本编辑成任意名人面孔
  • 适合需要高保真人脸动画的影视、虚拟人应用

生成式AI和文本到图像方法的快速发展重塑了我们对计算机生成图像的感知。与此同时,3D人脸重建领域也取得了显著进展,尤其基于3D可变形模型(3DMM)。本文提出SVP,一种新颖的2D/3D混合生成方法,利用大型预训练文本到图像先验(2D)生成逼真说话人脸视频,并通过3DMM进行控制(3D)。具体而言,我们对通用2D稳定扩散模型进行个性化微调,通过提供时间序列的3DMM作为条件并引入时序去噪过程,将其扩展为视频生成模型。输出为具有3DMM控制能力的个性化人物形象,其面部外观可按文本描述编辑为任意名人,无需测试时微调。该方法在定量与定性层面均进行了评估,结果表明其优于现有单目头部形象生成方法。

原文摘要 · Abstract (English)

Rapid advances in the field of generative AI and text-to-image methods in particular have transformed the way we interact with and perceive computer-generated imagery today. In parallel, much progress has been made in 3D face reconstruction, using 3D Morphable Models (3DMM). In this paper, we present SVP, a novel hybrid 2D/3D generation method that outputs photorealistic videos of talking faces leveraging a large pre-trained text-to-image prior (2D), controlled via a 3DMM (3D). Specifically, we introduce a person-specific fine-tuning of a general 2D stable diffusion model which we lift to a video model by providing temporal 3DMM sequences as conditioning and by introducing a temporal denoising procedure. As an output, this model generates temporally smooth imagery of a person with 3DMM-based controls, i.e., a person-specific avatar. The facial appearance of this person-specific avatar can be edited and morphed to text-defined celebrities, without any fine-tuning at test time. The method is analyzed quantitatively and qualitatively, and we show that our method outperforms state-of-the-art monocular head avatar methods.

人脸生成视频生成3DMM扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。