用文本控制生成可动的高保真3D人像,效果优于现有方法。
GaussianMotion: End-to-End Learning of Animatable Gaussian Avatars with Pose Guidance from Text
- 结合可变形高斯点云与文本到3D分数蒸馏,端到端训练动画人像。
- 在多种姿势下生成高质量纹理,静态和动态结果均领先基线。
- 适合需要高真实感、可自定义动作的虚拟角色生成场景。
本文提出GaussianMotion,一种基于高斯溅射的新型人体渲染模型,通过文本描述生成完全可动画化的3D人像。现有方法虽能实现文本到3D的人体生成,但在保真度、效率或姿态控制方面存在局限。本方法结合可变形3D高斯溅射与文本到3D分数蒸馏,端到端学习,实现任意姿态下的高保真渲染。通过优化阶段密集生成多样随机姿态,模型从姿态条件扩散模型中蒸馏出丰富自然运动特征。此外,提出自适应分数蒸馏策略,有效平衡细节真实感与表面平滑性。实验表明,该方法在静态与动态结果中均显著优于现有基线,能从不同文本输入生成多样化高质量3D人像。
原文摘要 · Abstract (English)
In this paper, we introduce GaussianMotion, a novel human rendering model that generates fully animatable scenes aligned with textual descriptions using Gaussian Splatting. Although existing methods achieve reasonable text-to-3D generation of human bodies using various 3D representations, they often face limitations in fidelity and efficiency, or primarily focus on static models with limited pose control. In contrast, our method generates fully animatable 3D avatars by combining deformable 3D Gaussian Splatting with text-to-3D score distillation, achieving high fidelity and efficient rendering for arbitrary poses. By densely generating diverse random poses during optimization, our deformable 3D human model learns to capture a wide range of natural motions distilled from a pose-conditioned diffusion model in an end-to-end manner. Furthermore, we propose Adaptive Score Distillation that effectively balances realistic detail and smoothness to achieve optimal 3D results. Experimental results demonstrate that our approach outperforms existing baselines by producing high-quality textures in both static and animated results, and by generating diverse 3D human models from various textual inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。