arXiv:2409.02851cs.CVcs.GR2024-09被引 7

用视频扩散模型生成单图3D人体,解决视角不一致问题。

Human-VDM: Learning Single-Image 3D Human Gaussian Splatting from Video Diffusion Models

  • 通过视频扩散模型生成连贯人体视频,保证视角一致性。
  • 经超分与插值增强后,3D高斯点云重建质量显著提升。
  • 适合需要高质量单图3D人体生成的研究与应用。

从单张RGB图像生成逼真的3D人体仍是计算机视觉中的挑战,需精确建模几何结构、高质量纹理及合理推测未见部分。现有方法多采用多视角扩散模型生成3D内容,但常因视角不一致影响生成质量。为此,我们提出Human-VDM,一种基于视频扩散模型的单图3D人体生成新方法。该方法利用高时序一致性视频生成3D人体,结合高斯点云(Gaussian Splatting)实现细节还原。整体由三模块构成:视图一致的人体视频扩散模块、视频增强模块(含超分辨率与插值)、以及3D高斯点云生成模块。首先输入单图至人体视频扩散模型生成连贯动作序列;随后通过超分和插值提升纹理清晰度与几何平滑性;最后在高质量一致视图引导下训练3D高斯点云模型。实验表明,Human-VDM在生成质量和数量上均超越当前最优方法。项目主页:https://human-vdm.github.io/Human-VDM/

原文摘要 · Abstract (English)

Generating lifelike 3D humans from a single RGB image remains a challenging task in computer vision, as it requires accurate modeling of geometry, high-quality texture, and plausible unseen parts. Existing methods typically use multi-view diffusion models for 3D generation, but they often face inconsistent view issues, which hinder high-quality 3D human generation. To address this, we propose Human-VDM, a novel method for generating 3D human from a single RGB image using Video Diffusion Models. Human-VDM provides temporally consistent views for 3D human generation using Gaussian Splatting. It consists of three modules: a view-consistent human video diffusion module, a video augmentation module, and a Gaussian Splatting module. First, a single image is fed into a human video diffusion module to generate a coherent human video. Next, the video augmentation module applies super-resolution and video interpolation to enhance the textures and geometric smoothness of the generated video. Finally, the 3D Human Gaussian Splatting module learns lifelike humans under the guidance of these high-resolution and view-consistent images. Experiments demonstrate that Human-VDM achieves high-quality 3D human from a single image, outperforming state-of-the-art methods in both generation quality and quantity. Project page: https://human-vdm.github.io/Human-VDM/

3D生成扩散模型高斯点云单图建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。