仅用一张图生成可文本控制的高保真3D虚拟人,无需后期处理。
Dream3DAvatar: Text-Controlled 3D Avatar Reconstruction from a Single Image
- 分两阶段:先生成多视角图像,再重建3D高斯点云。
- 引入姿态与身份适配器,精准控制肢体结构和面部特征。
- 支持文本编辑且对遮挡区域可控,适合虚拟偶像与游戏应用。
随着3D表示技术和生成模型的快速发展,仅凭单张图像重建全身3D虚拟人已取得显著进展。然而,由于单目输入信息有限,该任务本质上仍为病态问题,难以在生成过程中控制被遮挡区域的几何与纹理。为此,我们重构了重建流程,提出Dream3DAvatar——一种高效、可文本控制的两阶段3D虚拟人生成框架。第一阶段设计轻量级、适配器增强的多视角生成模型:引入Pose-Adapter注入SMPL-X渲染结果与骨骼信息,确保多视角几何与姿态一致性;通过ID-Adapter-G将高分辨率面部特征注入生成过程,保持面部身份一致;并利用BLIP2生成多视角图像的高质量文本描述,提升遮挡区域的文本可控性。第二阶段设计基于前馈Transformer的模型,配备多视角特征融合模块,从生成图像中重建高保真3D高斯溅射(3DGS)表示。此外,引入ID-Adapter-R,采用门控机制有效融合面部特征,提升高频细节恢复能力。大量实验表明,本方法可生成真实、可动画化的3D虚拟人,无需任何后处理,且在多个评估指标上持续优于现有基线。
原文摘要 · Abstract (English)
With the rapid advancement of 3D representation techniques and generative models, substantial progress has been made in reconstructing full-body 3D avatars from a single image. However, this task remains fundamentally ill-posedness due to the limited information available from monocular input, making it difficult to control the geometry and texture of occluded regions during generation. To address these challenges, we redesign the reconstruction pipeline and propose Dream3DAvatar, an efficient and text-controllable two-stage framework for 3D avatar generation. In the first stage, we develop a lightweight, adapter-enhanced multi-view generation model. Specifically, we introduce the Pose-Adapter to inject SMPL-X renderings and skeletal information into SDXL, enforcing geometric and pose consistency across views. To preserve facial identity, we incorporate ID-Adapter-G, which injects high-resolution facial features into the generation process. Additionally, we leverage BLIP2 to generate high-quality textual descriptions of the multi-view images, enhancing text-driven controllability in occluded regions. In the second stage, we design a feedforward Transformer model equipped with a multi-view feature fusion module to reconstruct high-fidelity 3D Gaussian Splat representations (3DGS) from the generated images. Furthermore, we introduce ID-Adapter-R, which utilizes a gating mechanism to effectively fuse facial features into the reconstruction process, improving high-frequency detail recovery. Extensive experiments demonstrate that our method can generate realistic, animation-ready 3D avatars without any post-processing and consistently outperforms existing baselines across multiple evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。