arXiv:2412.14963cs.CVcs.GR2024-12CVPR被引 68

单图秒生成高保真可动画3D人像,无需复杂后处理。

IDOL: Instant Photorealistic 3D Human Creation from a Single Image

  • 用10万张多视角人像构建新数据集,支持多样姿态与外观
  • 单图输入即可在1秒内重建1024分辨率的逼真人像
  • 生成结果可直接动画化,还支持形状与纹理编辑

从单张图像高保真、可动画地重建全身3D人像是一项挑战,因人体外观与姿态多样且高质量训练数据稀缺。本文从数据、模型和表征三方面重新思考该任务:首先构建大规模人像中心生成数据集HuGe100K,包含10万组由可控姿态图像生成的24视角图像;其次基于该数据集,设计可扩展的前馈式Transformer模型,从单张图像预测统一空间中的3D人形高斯表示,实现姿态、体型、服装几何与纹理的解耦;所估计的高斯可直接用于动画而无需后处理。实验验证了方法的有效性,模型仅用单个GPU即可在1秒内完成1024分辨率的逼真人像重建,并支持多种应用及形状、纹理编辑任务。

原文摘要 · Abstract (English)

Creating a high-fidelity, animatable 3D full-body avatar from a single image is a challenging task due to the diverse appearance and poses of humans and the limited availability of high-quality training data. To achieve fast and high-quality human reconstruction, this work rethinks the task from the perspectives of dataset, model, and representation. First, we introduce a large-scale HUman-centric GEnerated dataset, HuGe100K, consisting of 100K diverse, photorealistic sets of human images. Each set contains 24-view frames in specific human poses, generated using a pose-controllable image-to-multi-view model. Next, leveraging the diversity in views, poses, and appearances within HuGe100K, we develop a scalable feed-forward transformer model to predict a 3D human Gaussian representation in a uniform space from a given human image. This model is trained to disentangle human pose, body shape, clothing geometry, and texture. The estimated Gaussians can be animated without post-processing. We conduct comprehensive experiments to validate the effectiveness of the proposed dataset and method. Our model demonstrates the ability to efficiently reconstruct photorealistic humans at 1K resolution from a single input image using a single GPU instantly. Additionally, it seamlessly supports various applications, as well as shape and texture editing tasks. Project page: https://yiyuzhuang.github.io/IDOL/.

3D人像单图生成高斯表示实时重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。