用生成模型先验重建单图3D虚拟人,更真实且可动画
MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar Reconstruction
- 用生成式3D先验引导单图重建,保持几何与外观一致性
- 在真实场景中表现优于现有方法,生成结果无模糊与失真
- 输出可直接动画化,适合虚拟人生成与数字孪生应用
我们提出MoGA,一种从单张图像重建高保真3D Gaussian虚拟人的新方法。核心挑战在于推断未见视角的外观与几何细节,同时保证3D一致性和真实性。以往方法依赖2D扩散模型生成未见视图,但生成结果稀疏且不一致,导致3D伪影和模糊外观。为此,我们引入一个生成式虚拟人模型,通过从学习到的先验分布中采样变形的Gaussians来生成多样化的3D虚拟人。由于3D训练数据有限,该模型难以捕捉所有未见身份的细节,因此我们将其作为先验,将输入图像投影至其隐空间,并施加额外的3D外观与几何约束以确保一致性。我们的方法将Gaussian虚拟人构建视为模型反演,通过拟合生成式虚拟人与2D扩散模型合成视图。生成模型提供初始化、3D正则化并辅助姿态优化。实验表明,本方法超越当前最优技术,且在真实场景下具有良好泛化能力。生成的Gaussian虚拟人本身具备天然可动画性。
原文摘要 · Abstract (English)
We present MoGA, a novel method to reconstruct high-fidelity 3D Gaussian avatars from a single-view image. The main challenge lies in inferring unseen appearance and geometric details while ensuring 3D consistency and realism. Most previous methods rely on 2D diffusion models to synthesize unseen views; however, these generated views are sparse and inconsistent, resulting in unrealistic 3D artifacts and blurred appearance. To address these limitations, we leverage a generative avatar model, that can generate diverse 3D avatars by sampling deformed Gaussians from a learned prior distribution. Due to limited 3D training data, such a 3D model alone cannot capture all image details of unseen identities. Consequently, we integrate it as a prior, ensuring 3D consistency by projecting input images into its latent space and enforcing additional 3D appearance and geometric constraints. Our novel approach formulates Gaussian avatar creation as model inversion by fitting the generative avatar to synthetic views from 2D diffusion models. The generative avatar provides an initialization for model fitting, enforces 3D regularization, and helps in refining pose. Experiments show that our method surpasses state-of-the-art techniques and generalizes well to real-world scenarios. Our Gaussian avatars are also inherently animatable. For code, see https://zj-dong.github.io/MoGA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。