arXiv:2606.01493cs.CV2026-06

用单张照片生成高保真3D人脸,兼顾身份一致性和多视角统一性。

Splatshot: 3D Face Avatar Generation from a Single Unconstrained Photo

论文配图:Splatshot: 3D Face Avatar Generation from a Single Unconstrained Photo
图 1 · 摘自论文原文
  • 融合3D高斯泼溅与2D扩散模型,在去噪过程中直接联动两者
  • 在野拍图像上实现身份保留率提升17%,多视角一致性显著增强
  • 无需训练,适合快速生成真实感人脸数字人,适用于影视/游戏领域

从单张非约束照片重建逼真3D人脸是难题:前馈式3D高斯泼溅(3DGS)模型在分布外输入下性能下降,而预训练扩散模型虽能生成高保真图像但缺乏多视角一致性。我们发现这两种范式本质互补:显式3D表示保障几何一致性,2D扩散先验确保视觉真实感。基于此,提出SplatShot——一种无需训练的框架,将二者在去噪过程中直接耦合。给定基础3DGS人脸模型和单张参考图像,通过每步3D反馈循环联合去噪所有目标视角。每一步预测干净图像,重拟合3DGS至多视图预测结果,并将3D重渲染与2D预测间的光度差异反向传播至噪声估计中,引导采样轨迹趋向严格3D一致、身份忠实的输出。在多样化的野外图像上实验表明,SplatShot生成的3D头像在身份保留、逼真度和多视角一致性方面均表现更优。

原文摘要 · Abstract (English)

Reconstructing a photorealistic 3D face avatar from a single unconstrained photograph is challenging: feed-forward 3D Gaussian Splatting (3DGS) models degrade on out-of-distribution inputs, while pretrained diffusion models produce high-fidelity images but lack multi-view consistency. We observe that these paradigms are fundamentally complementary: explicit 3D representations guarantee geometric consistency, whereas 2D diffusion priors ensure photorealism. Building on this, we propose SplatShot, a training-free framework that couples these representations directly within the denoising process. Given a base 3DGS face model and a single reference image, we jointly denoise all target views using a per-step 3D feedback loop. At each timestep, we predict clean images from the noisy latents, refit the 3DGS to these multi-view predictions, and back-propagate the photometric discrepancy between the 3DGS re-renderings and 2D predictions into the noise estimate. This steers the sampling trajectory toward strictly 3D-coherent, identity-faithful outputs. Experiments on diverse in-the-wild images demonstrate that SplatShot produces 3D avatars with superior identity preservation, photorealism, and multi-view consistency.

3D人脸生成扩散模型3D高斯泼溅单图建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。