arXiv:2603.04307cs.CV2026-03

用双扩散模型10秒生成高保真3D虚拟人,支持图文联合控制。

Dual Diffusion Models for Multi-modal Guided 3D Avatar Generation

  • 构建双扩散架构:纹理与几何模型分别响应图文提示。
  • 10万级多模态数据训练,生成速度<10秒,无迭代优化。
  • 适合需要快速生成精细3D角色的VR/交互应用开发者。

从文本或图像提示生成高保真3D虚拟人是虚拟现实和人机交互中的重要需求。现有文本驱动方法依赖迭代的Score Distillation Sampling(SDS)或CLIP优化,难以实现细粒度语义控制且推理过慢;而图像驱动方法受限于高质量3D人脸扫描数据稀缺及获取成本高,制约模型泛化能力。为此,我们构建了一个包含超过10万对样本的大型多模态数据集,涵盖细粒度文本描述、真实场景人脸图像、高质量光照归一化纹理UV图及3D几何形状。基于此数据集,提出PromptAvatar框架,采用双扩散模型:纹理扩散模型(TDM)支持文本和/或图像的多条件引导,几何扩散模型(GDM)由文本提示引导。通过直接学习多模态提示到3D表示的映射,该方法无需耗时的迭代优化,可在10秒内生成无阴影、高保真的3D虚拟人。大量定量与定性实验表明,本方法在生成质量、细节对齐精度及计算效率上显著优于现有最先进方法。

原文摘要 · Abstract (English)

Generating high-fidelity 3D avatars from text or image prompts is highly sought after in virtual reality and human-computer interaction. However, existing text-driven methods often rely on iterative Score Distillation Sampling (SDS) or CLIP optimization, which struggle with fine-grained semantic control and suffer from excessively slow inference. Meanwhile, image-driven approaches are severely bottlenecked by the scarcity and high acquisition cost of high-quality 3D facial scans, limiting model generalization. To address these challenges, we first construct a novel, large-scale dataset comprising over 100,000 pairs across four modalities: fine-grained textual descriptions, in-the-wild face images, high-quality light-normalized texture UV maps, and 3D geometric shapes. Leveraging this comprehensive dataset, we propose PromptAvatar, a framework featuring dual diffusion models. Specifically, it integrates a Texture Diffusion Model (TDM) that supports flexible multi-condition guidance from text and/or image prompts, alongside a Geometry Diffusion Model (GDM) guided by text prompts. By learning the direct mapping from multi-modal prompts to 3D representations, PromptAvatar eliminates the need for time-consuming iterative optimization, successfully generating high-fidelity, shading-free 3D avatars in under 10 seconds. Extensive quantitative and qualitative experiments demonstrate that our method significantly outperforms existing state-of-the-art approaches in generation quality, fine-grained detail alignment, and computational efficiency.

3D生成扩散模型多模态虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。