arXiv:2502.06957cs.CV2025-02ICCV被引 18

仅用一张图生成逼真连贯的3D虚拟形象,解决多视角不一致问题。

GAS: Generative Avatar Synthesis from a Single Image

  • 结合回归重建与扩散模型,用NeRF提供完整条件信息
  • 在真实场景数据上实现跨域泛化,多视角和时间一致性好
  • 适合需要高质量3D虚拟角色的应用,如影视或元宇宙

我们提出一种统一且通用的框架,仅需单张图像即可生成视图一致且时序连贯的3D虚拟形象,解决了单图生成虚拟形象这一难题。现有基于扩散模型的方法通常依赖稀疏人体模板(如深度图或法向图),因信号与真实外观不匹配,导致多视角和时序不一致。我们的方法通过融合回归式3D人体重建的重构能力与扩散模型的生成能力,首先利用广义NeRF生成初始3D人体,提供全面的条件信息,确保合成结果忠实于参考外观与结构;随后,将该NeRF推导出的几何与外观特征输入视频扩散模型,战略性地保障生成过程中多视角与时间的一致性。实验证明,所提方法在多种域内与域外真实场景数据集上均表现出优异的泛化能力。

原文摘要 · Abstract (English)

We present a unified and generalizable framework for synthesizing view-consistent and temporally coherent avatars from a single image, addressing the challenging task of single-image avatar generation. Existing diffusion-based methods often condition on sparse human templates (e.g., depth or normal maps), which leads to multi-view and temporal inconsistencies due to the mismatch between these signals and the true appearance of the subject. Our approach bridges this gap by combining the reconstruction power of regression-based 3D human reconstruction with the generative capabilities of a diffusion model. In a first step, an initial 3D reconstructed human through a generalized NeRF provides comprehensive conditioning, ensuring high-quality synthesis faithful to the reference appearance and structure. Subsequently, the derived geometry and appearance from the generalized NeRF serve as input to a video-based diffusion model. This strategic integration is pivotal for enforcing both multi-view and temporal consistency throughout the avatar's generation. Empirical results underscore the superior generalization ability of our proposed method, demonstrating its effectiveness across diverse in-domain and out-of-domain in-the-wild datasets.

3D生成虚拟形象扩散模型单图生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。