arXiv:2510.23929cs.CV2025-10

用单步扩散模型快速生成高保真人脸新视角图像。

TurboPortrait3D: Single-step diffusion-based fast portrait novel-view synthesis

  • 结合扩散模型与3D感知,单步优化多视角一致性
  • 在真实人脸数据上训练,生成图像细节更丰富
  • 仅需一张正面照,推理速度快,适合实时应用

我们提出TurboPortrait3D,一种低延迟的人脸新视角合成方法。现有图像到3D模型虽能生成可渲染的3D表示,但常伴随视觉伪影、细节不足且难以保留身份特征。而图像扩散模型虽能生成高质量图像,却因缺乏3D基础,难以保证多视角一致性。本工作证明,可通过图像空间扩散模型显著提升现有图像到虚拟角色方法的质量,同时保持3D感知并实现低延迟。输入仅需一张人物正面图像,通过前馈式图像到虚拟角色生成流程获得初始3D表示及噪声渲染图,再将这些噪声图输入一个单步扩散模型,该模型以输入图像为条件,并专门训练用于多视角一致地细化渲染结果。此外,我们引入一种新训练策略:先在大规模合成多视角数据上预训练,再在高质量真实图像上微调。实验表明,该方法在定性和定量上均优于当前最先进水平,且时间效率高。

原文摘要 · Abstract (English)

We introduce TurboPortrait3D: a method for low-latency novel-view synthesis of human portraits. Our approach builds on the observation that existing image-to-3D models for portrait generation, while capable of producing renderable 3D representations, are prone to visual artifacts, often lack of detail, and tend to fail at fully preserving the identity of the subject. On the other hand, image diffusion models excel at generating high-quality images, but besides being computationally expensive, are not grounded in 3D and thus are not directly capable of producing multi-view consistent outputs. In this work, we demonstrate that image-space diffusion models can be used to significantly enhance the quality of existing image-to-avatar methods, while maintaining 3D-awareness and running with low-latency. Our method takes a single frontal image of a subject as input, and applies a feedforward image-to-avatar generation pipeline to obtain an initial 3D representation and corresponding noisy renders. These noisy renders are then fed to a single-step diffusion model which is conditioned on input image(s), and is specifically trained to refine the renders in a multi-view consistent way. Moreover, we introduce a novel effective training strategy that includes pre-training on a large corpus of synthetic multi-view data, followed by fine-tuning on high-quality real images. We demonstrate that our approach both qualitatively and quantitatively outperforms current state-of-the-art for portrait novel-view synthesis, while being efficient in time.

3D人脸扩散模型新视角合成实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。