TeRA用两阶段训练生成逼真3D虚拟人,支持文本控制局部修改。
TeRA: Rethinking Text-guided Realistic 3D Avatar Generation
- 先蒸馏大模型得结构化潜空间,再用文本控制扩散模型生成
- 无需迭代优化,生成速度更快,主观与客观评价均更优
- 适合需要快速定制真实感3D角色的创作者和游戏开发者
本文提出TeRA,一种更高效、更有效的文本引导3D虚拟人生成框架,优于以往基于SDS的模型及通用大3D生成模型。该方法采用两阶段训练策略,首先从大型人体重建模型中蒸馏出解码器,构建结构化的潜空间;随后在该潜空间内训练一个文本可控的潜扩散模型,以生成逼真3D人类虚拟人。TeRA通过消除缓慢的迭代优化过程提升模型性能,并实现基于结构化3D人体表示的文本驱动局部定制。实验表明,该方法在主观与客观评估中均优于现有文本到虚拟人生成模型。
原文摘要 · Abstract (English)
In this paper, we rethink text-to-avatar generative models by proposing TeRA, a more efficient and effective framework than the previous SDS-based models and general large 3D generative models. Our approach employs a two-stage training strategy for learning a native 3D avatar generative model. Initially, we distill a decoder to derive a structured latent space from a large human reconstruction model. Subsequently, a text-controlled latent diffusion model is trained to generate photorealistic 3D human avatars within this latent space. TeRA enhances the model performance by eliminating slow iterative optimization and enables text-based partial customization through a structured 3D human representation. Experiments have proven our approach's superiority over previous text-to-avatar generative models in subjective and objective evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。