用百万级3D人体高斯资产,实现快速高质量人体建模。
SIGMAN:Scaling 3D Human Gaussian Generation with Millions of Assets
- 通过基于VAE的压缩与DiT生成,将多视角图像转为高斯表示。
- 在100万规模数据集上训练,还原出纹理精细、衣物动态逼真的3D人体。
- 适合需要大规模高质量3D人体生成的虚拟人、游戏、影视应用。
3D人体数字化长期面临挑战。现有方法多基于优化或前馈框架(单视图回归或多视图生成),受限于速度慢、质量低、级联推理及遮挡导致的低维到高维映射模糊。同时,3D人体数据集规模小,难以支撑大模型训练。为此,我们提出一种潜空间生成范式:先用结构化UV的变分自编码器(VAE)将多视角图像压缩为高斯表示,再结合基于DiT的条件生成,将原本病态的低维到高维映射转化为可学习的分布迁移,支持端到端推理。此外,结合多视图优化与合成数据构建了包含100万3D高斯资产的HGS-1M数据集,支持大规模训练。实验表明,该范式在大规模训练下生成了具有复杂纹理、面部细节和松散衣物变形的高质量3D人体高斯模型。
原文摘要 · Abstract (English)
3D human digitization has long been a highly pursued yet challenging task. Existing methods aim to generate high-quality 3D digital humans from single or multiple views, but remain primarily constrained by current paradigms and the scarcity of 3D human assets. Specifically, recent approaches fall into several paradigms: optimization-based and feed-forward (both single-view regression and multi-view generation with reconstruction). However, they are limited by slow speed, low quality, cascade reasoning, and ambiguity in mapping low-dimensional planes to high-dimensional space due to occlusion and invisibility, respectively. Furthermore, existing 3D human assets remain small-scale, insufficient for large-scale training. To address these challenges, we propose a latent space generation paradigm for 3D human digitization, which involves compressing multi-view images into Gaussians via a UV-structured VAE, along with DiT-based conditional generation, we transform the ill-posed low-to-high-dimensional mapping problem into a learnable distribution shift, which also supports end-to-end inference. In addition, we employ the multi-view optimization approach combined with synthetic data to construct the HGS-1M dataset, which contains $1$ million 3D Gaussian assets to support the large-scale training. Experimental results demonstrate that our paradigm, powered by large-scale training, produces high-quality 3D human Gaussians with intricate textures, facial details, and loose clothing deformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。