用骨骼控制提升2D扩散模型生成高质量可动3D人像
DreamWaltz-G: Expressive 3D Gaussian Avatars from Skeleton-Guided 2D Diffusion
- 将3D人体骨架引导融入2D扩散模型,增强姿态与视角一致性
- 生成的3D人像无多脸、多余肢体,且支持实时渲染与表情动画
- 适合需要高保真可动数字人的影视/游戏/虚拟主播应用
利用预训练的2D扩散模型和得分蒸馏采样(SDS),近期方法在文本到3D人像生成方面取得良好进展。然而,生成具备表现力动画能力的高质量3D人像仍具挑战。本文提出DreamWaltz-G,一种从文本生成可动画3D人像的新框架。其核心为骨骼引导得分蒸馏与混合3D高斯人像表示。具体而言,骨骼引导得分蒸馏将3D人体模板的骨架控制融入2D扩散模型,提升了在视角与人体姿态上的SDS监督一致性,有效缓解了多脸、多余肢体、模糊等问题。混合3D高斯人像表示基于高效3D高斯,结合神经隐式场与参数化3D网格,实现实时渲染、稳定SDS优化与丰富动画表达。大量实验表明,DreamWaltz-G在视觉质量与动画表现力上均优于现有方法。该框架还可用于人体视频重演与多主体场景构建等多样化应用。
原文摘要 · Abstract (English)
Leveraging pretrained 2D diffusion models and score distillation sampling (SDS), recent methods have shown promising results for text-to-3D avatar generation. However, generating high-quality 3D avatars capable of expressive animation remains challenging. In this work, we present DreamWaltz-G, a novel learning framework for animatable 3D avatar generation from text. The core of this framework lies in Skeleton-guided Score Distillation and Hybrid 3D Gaussian Avatar representation. Specifically, the proposed skeleton-guided score distillation integrates skeleton controls from 3D human templates into 2D diffusion models, enhancing the consistency of SDS supervision in terms of view and human pose. This facilitates the generation of high-quality avatars, mitigating issues such as multiple faces, extra limbs, and blurring. The proposed hybrid 3D Gaussian avatar representation builds on the efficient 3D Gaussians, combining neural implicit fields and parameterized 3D meshes to enable real-time rendering, stable SDS optimization, and expressive animation. Extensive experiments demonstrate that DreamWaltz-G is highly effective in generating and animating 3D avatars, outperforming existing methods in both visual quality and animation expressiveness. Our framework further supports diverse applications, including human video reenactment and multi-subject scene composition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。