仅用一张图生成逼真连贯的3D人像,靠扩散模型补全缺失信息
HumanGif: Single-View Human Diffusion with Generative Prior
- 用扩散模型结合先验知识,从单张图生成3D人像
- 在多个数据集上达到最佳视觉效果,泛化能力强
- 适合做虚拟人、动画角色生成的研究者和开发者
以往的3D人体生成方法在稀疏视图图像或单目视频基础上已能生成视角一致且时间连贯的结果。然而,在仅提供单张图像的情况下,仍难以生成持续逼真、视角一致且时间连贯的人体虚拟形象,因输入信息有限。受2D角色动画成功的启发,我们提出HumanGif,一种基于生成先验的单视图人体扩散模型。具体而言,将单视图3D人体新视角与姿态合成建模为单视图条件下的扩散过程,利用基础扩散模型的生成先验来补充缺失信息。为确保新视角与姿态生成的精细一致性,我们在HumanGif中引入了人类NeRF模块,从输入图像中学习空间对齐特征,隐式捕捉相机与人体姿态间的相对变换。此外,在优化过程中引入图像级损失,以弥合扩散模型中潜在空间与图像空间之间的差距。在RenderPeople、DNA-Rendering、THuman 2.1和TikTok数据集上的大量实验表明,HumanGif在感知质量上表现最佳,且在新视角与新姿态合成方面具有更强的泛化能力。
原文摘要 · Abstract (English)
Previous 3D human creation methods have made significant progress in synthesizing view-consistent and temporally aligned results from sparse-view images or monocular videos. However, it remains challenging to produce perpetually realistic, view-consistent, and temporally coherent human avatars from a single image, as limited information is available in the single-view input setting. Motivated by the success of 2D character animation, we propose HumanGif, a single-view human diffusion model with generative prior. Specifically, we formulate the single-view-based 3D human novel view and pose synthesis as a single-view-conditioned human diffusion process, utilizing generative priors from foundational diffusion models to complement the missing information. To ensure fine-grained and consistent novel view and pose synthesis, we introduce a Human NeRF module in HumanGif to learn spatially aligned features from the input image, implicitly capturing the relative camera and human pose transformation. Furthermore, we introduce an image-level loss during optimization to bridge the gap between latent and image spaces in diffusion models. Extensive experiments on RenderPeople, DNA-Rendering, THuman 2.1, and TikTok datasets demonstrate that HumanGif achieves the best perceptual performance, with better generalizability for novel view and pose synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。