用文本生成可动画的3D头像,解决外观与模型对不齐的问题。
Text-based Animatable 3D Avatars with Morphable Model Alignment
- 用预训练模型初始化3D头像,提升外观和结构稳定性。
- 引入控制网络,基于形态模型的语义与法线图优化表情动态。
- 生成结果更真实,对齐更准确,适合虚拟角色创作。
从文本生成高质量、可动画的3D头部虚拟人,在游戏、影视和虚拟助手等领域有巨大潜力。现有方法通常结合参数化头模与2D扩散模型,通过得分蒸馏采样生成3D一致结果,但难以还原真实细节,且常因外观与驱动参数模型错位导致动画不自然。我们发现其根源在于2D扩散预测在3D蒸馏过程中存在歧义:一是文本输入对头像外观与几何约束不足;二是扩散模型无法融入参数模型信息,导致语义对齐不足。为此,我们提出AnimPortrait3D框架,引入两种策略:首先利用预训练文本到3D模型提供先验,初始化具备稳健外观、几何与绑定关系的3D头像;其次使用条件控制网络,基于形态模型的语义图与法线图细化动态表情,确保精确对齐。实验表明,该方法在合成质量、对齐度与动画保真度上均超越现有技术,推动了文本驱动可动画3D头像生成的最新进展。
原文摘要 · Abstract (English)
The generation of high-quality, animatable 3D head avatars from text has enormous potential in content creation applications such as games, movies, and embodied virtual assistants. Current text-to-3D generation methods typically combine parametric head models with 2D diffusion models using score distillation sampling to produce 3D-consistent results. However, they struggle to synthesize realistic details and suffer from misalignments between the appearance and the driving parametric model, resulting in unnatural animation results. We discovered that these limitations stem from ambiguities in the 2D diffusion predictions during 3D avatar distillation, specifically: i) the avatar's appearance and geometry is underconstrained by the text input, and ii) the semantic alignment between the predictions and the parametric head model is insufficient because the diffusion model alone cannot incorporate information from the parametric model. In this work, we propose a novel framework, AnimPortrait3D, for text-based realistic animatable 3DGS avatar generation with morphable model alignment, and introduce two key strategies to address these challenges. First, we tackle appearance and geometry ambiguities by utilizing prior information from a pretrained text-to-3D model to initialize a 3D avatar with robust appearance, geometry, and rigging relationships to the morphable model. Second, we refine the initial 3D avatar for dynamic expressions using a ControlNet that is conditioned on semantic and normal maps of the morphable model to ensure accurate alignment. As a result, our method outperforms existing approaches in terms of synthesis quality, alignment, and animation fidelity. Our experiments show that the proposed method advances the state of the art in text-based, animatable 3D head avatar generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。