仅用一张照片生成高质量360度三维人脸,通用性强。
FaceLift: Learning Generalizable Single Image 3D Face Reconstruction from Synthetic Heads
- 用多视角扩散模型生成背面和侧视图,再用Transformer重建3D高斯点云
- 在合成数据上训练,仍能精准还原真实照片中的人脸身份与细节
- 适合需要高保真三维人脸重建的虚拟人、数字孪生场景
我们提出FaceLift,一种无需迭代的前馈方法,仅凭单张图像即可实现高质量、全视角360度三维头部重建。该流程首先通过多视角潜在扩散模型,从单一面部输入生成一致的侧面和背面视图,随后输入基于Transformer的重建器,生成完整的3D高斯点云表示。以往单目三维人脸重建方法常因缺乏多视角监督而存在视角覆盖不全或视图不一致的问题。为此,我们构建了一个高质量的合成头部数据集,实现跨视角的一致性监督。为缩小合成数据与真实图像之间的域差距,我们提出一种简单有效的技术:在生成新视角的同时学习重建输入图像,从而保持生成过程对输入的忠实度。尽管仅在合成数据上训练,本方法在真实图像上仍表现出卓越的泛化能力。通过大量定性和定量评估,结果表明FaceLift在身份保留、细节恢复和渲染质量方面均优于当前最先进方法。
原文摘要 · Abstract (English)
We present FaceLift, a novel feed-forward approach for generalizable high-quality 360-degree 3D head reconstruction from a single image. Our pipeline first employs a multi-view latent diffusion model to generate consistent side and back views from a single facial input, which then feeds into a transformer-based reconstructor that produces a comprehensive 3D Gaussian splats representation. Previous methods for monocular 3D face reconstruction often lack full view coverage or view consistency due to insufficient multi-view supervision. We address this by creating a high-quality synthetic head dataset that enables consistent supervision across viewpoints. To bridge the domain gap between synthetic training data and real-world images, we propose a simple yet effective technique that ensures the view generation process maintains fidelity to the input by learning to reconstruct the input image alongside the view generation. Despite being trained exclusively on synthetic data, our method demonstrates remarkable generalization to real-world images. Through extensive qualitative and quantitative evaluations, we show that FaceLift outperforms state-of-the-art 3D face reconstruction methods on identity preservation, detail recovery, and rendering quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。