arXiv:2412.02684cs.CVcs.AI2024-12CVPR被引 42

单图生成可动画真人虚拟形象,实时渲染且视角一致。

AniGS: Animatable Gaussian Avatar from a Single Image with Inconsistent Gaussian Reconstruction

  • 用生成模型先产多视角标准姿势图像,辅助3D重建
  • 采用4D高斯点云实现不一致视角下的实时渲染
  • 适合数字人、虚拟主播等需快速生成的场景

从单张图像生成可动画的人类虚拟形象对数字人建模至关重要。现有3D重建方法难以捕捉精细细节,而生成式动画方法虽避免显式3D建模,却在极端姿态下存在视角不一致和计算效率低的问题。本文利用生成模型生成多视角标准姿态图像及法线图,解决可动画人体重建中的歧义问题。通过将重建任务重构为4D问题,提出基于4D高斯溅射的高效3D建模方法,支持推理时实时渲染。我们采用基于Transformer的视频生成模型,在大规模视频数据集上预训练以提升泛化能力。实验表明,该方法能从真实环境图像中生成逼真、实时动画的3D人体形象,验证了其有效性和泛化能力。

原文摘要 · Abstract (English)

Generating animatable human avatars from a single image is essential for various digital human modeling applications. Existing 3D reconstruction methods often struggle to capture fine details in animatable models, while generative approaches for controllable animation, though avoiding explicit 3D modeling, suffer from viewpoint inconsistencies in extreme poses and computational inefficiencies. In this paper, we address these challenges by leveraging the power of generative models to produce detailed multi-view canonical pose images, which help resolve ambiguities in animatable human reconstruction. We then propose a robust method for 3D reconstruction of inconsistent images, enabling real-time rendering during inference. Specifically, we adapt a transformer-based video generation model to generate multi-view canonical pose images and normal maps, pretraining on a large-scale video dataset to improve generalization. To handle view inconsistencies, we recast the reconstruction problem as a 4D task and introduce an efficient 3D modeling approach using 4D Gaussian Splatting. Experiments demonstrate that our method achieves photorealistic, real-time animation of 3D human avatars from in-the-wild images, showcasing its effectiveness and generalization capability.

3D生成可动画虚拟人高斯溅射单图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。