arXiv:2501.15008cs.CV2025-01被引 1

仅用一张图就能生成可任意视角观看的3D人体模型

HuGDiffusion: Generalizable Single-Image Human Rendering via 3D Gaussian Diffusion

  • 用扩散模型从单张图学习人体3D高斯属性
  • 多阶段生成策略提升3DGS属性还原精度
  • 适合需要快速建模真实人物的影视/游戏开发者

我们提出HuGDiffusion,一种可泛化的3D高斯泼溅(3DGS)学习框架,仅需单张图像即可实现人物角色的新视角合成(NVS)。现有方法通常依赖单目视频或已标定的多视角图像输入,在相机姿态未知或随意的真实场景中应用受限。本文通过扩散模型,结合从单图提取的人体先验信息,生成3DGS属性集合。基于实证观察,联合优化全部3DGS属性难以收敛,因此设计多阶段生成策略分别获取不同类型的3DGS属性。为促进训练,我们构建了代理真值3D高斯属性作为高质量属性级监督信号。大量实验表明,HuGDiffusion在性能上显著优于当前最先进方法。代码将公开。

原文摘要 · Abstract (English)

We present HuGDiffusion, a generalizable 3D Gaussian splatting (3DGS) learning pipeline to achieve novel view synthesis (NVS) of human characters from single-view input images. Existing approaches typically require monocular videos or calibrated multi-view images as inputs, whose applicability could be weakened in real-world scenarios with arbitrary and/or unknown camera poses. In this paper, we aim to generate the set of 3DGS attributes via a diffusion-based framework conditioned on human priors extracted from a single image. Specifically, we begin with carefully integrated human-centric feature extraction procedures to deduce informative conditioning signals. Based on our empirical observations that jointly learning the whole 3DGS attributes is challenging to optimize, we design a multi-stage generation strategy to obtain different types of 3DGS attributes. To facilitate the training process, we investigate constructing proxy ground-truth 3D Gaussian attributes as high-quality attribute-level supervision signals. Through extensive experiments, our HuGDiffusion shows significant performance improvements over the state-of-the-art methods. Our code will be made publicly available.

3D建模扩散模型单图生成人体渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。