arXiv:2411.18624cs.CV2024-11NeurIPS被引 6

无需参数化人体模型,单图生成高质量3D人体模型。

GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data

  • 构建2D与3D扩散先验,不依赖SMPL等参数模型。
  • 在真实图像上实现高保真3D几何与纹理重建。
  • 对自然姿态和不同体型具有强泛化能力,适合实际应用。

给定一张野外拍摄的人体照片,重建高保真3D人体模型仍是难题。现有方法面临身体比例多变、个人物品多样、姿态模糊与纹理不一致等问题,且高质量人体数据稀缺。为此,我们提出通用的图像到3D人体重建框架GeneMAN,基于包含3D扫描、多视角视频、单图及自动生成合成数据的多源高质量人体数据集。GeneMAN包含三个核心模块:1)训练人体专属文本到图像扩散模型与视图条件扩散模型,分别作为2D与3D人体先验;2)利用预训练先验,通过几何初始化与精修流程恢复高质量3D人体几何;3)通过多空间纹理精修流程,在隐空间与像素空间连续优化纹理。大量实验表明,GeneMAN可从单张图像生成高质量3D人体模型,优于现有最先进方法。尤其在处理野外图像时展现更强泛化能力,能生成自然姿态下带有常见物品的高质量3D模型,不受输入图像中身体比例影响。

原文摘要 · Abstract (English)

Given a single in-the-wild human photo, it remains a challenging task to reconstruct a high-fidelity 3D human model. Existing methods face difficulties including a) the varying body proportions captured by in-the-wild human images; b) diverse personal belongings within the shot; and c) ambiguities in human postures and inconsistency in human textures. In addition, the scarcity of high-quality human data intensifies the challenge. To address these problems, we propose a Generalizable image-to-3D huMAN reconstruction framework, dubbed GeneMAN, building upon a comprehensive multi-source collection of high-quality human data, including 3D scans, multi-view videos, single photos, and our generated synthetic human data. GeneMAN encompasses three key modules. 1) Without relying on parametric human models (e.g., SMPL), GeneMAN first trains a human-specific text-to-image diffusion model and a view-conditioned diffusion model, serving as GeneMAN 2D human prior and 3D human prior for reconstruction, respectively. 2) With the help of the pretrained human prior models, the Geometry Initialization-&-Sculpting pipeline is leveraged to recover high-quality 3D human geometry given a single image. 3) To achieve high-fidelity 3D human textures, GeneMAN employs the Multi-Space Texture Refinement pipeline, consecutively refining textures in the latent and the pixel spaces. Extensive experimental results demonstrate that GeneMAN could generate high-quality 3D human models from a single image input, outperforming prior state-of-the-art methods. Notably, GeneMAN could reveal much better generalizability in dealing with in-the-wild images, often yielding high-quality 3D human models in natural poses with common items, regardless of the body proportions in the input images.

3D重建扩散模型人体建模单图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。