用单目视频重建可驱动的逼真3D人脸,解决视角缺失问题。
GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view Diffusion
- 通过多视角扩散模型填补未观测区域,保持渲染一致性。
- 在NeRSemble数据集上新视角合成效果优于现有方法。
- 适合手机拍摄视频重建高保真人脸,对身份细节还原精准。
我们提出一种新方法,从智能手机等普通设备拍摄的单目视频中重建可驱动的3D高斯人脸。由于观测受限,未观测区域易产生伪影,影响新视角一致性。为此,我们引入多视角头部扩散模型,利用其先验填补缺失区域并确保高斯点云渲染的一致性。为实现精确视角控制,采用基于FLAME的头像重建生成法线图,提供像素级归纳偏置;同时以输入图像的VAE特征作为条件,保留面部身份与外观细节。在高斯人脸重建中,通过迭代去噪图像作为伪真值,蒸馏多视角扩散先验,有效缓解过饱和问题。为进一步提升逼真度,应用潜在空间上采样先验,在解码前优化去噪潜在表示。我们在NeRSemble数据集上评估,结果表明GAF在新视角合成上超越先前最优方法,并实现了消费级设备拍摄视频的更高保真度重建。
原文摘要 · Abstract (English)
We propose a novel approach for reconstructing animatable 3D Gaussian avatars from monocular videos captured by commodity devices like smartphones. Photorealistic 3D head avatar reconstruction from such recordings is challenging due to limited observations, which leaves unobserved regions under-constrained and can lead to artifacts in novel views. To address this problem, we introduce a multi-view head diffusion model, leveraging its priors to fill in missing regions and ensure view consistency in Gaussian splatting renderings. To enable precise viewpoint control, we use normal maps rendered from FLAME-based head reconstruction, which provides pixel-aligned inductive biases. We also condition the diffusion model on VAE features extracted from the input image to preserve facial identity and appearance details. For Gaussian avatar reconstruction, we distill multi-view diffusion priors by using iteratively denoised images as pseudo-ground truths, effectively mitigating over-saturation issues. To further improve photorealism, we apply latent upsampling priors to refine the denoised latent before decoding it into an image. We evaluate our method on the NeRSemble dataset, showing that GAF outperforms previous state-of-the-art methods in novel view synthesis. Furthermore, we demonstrate higher-fidelity avatar reconstructions from monocular videos captured on commodity devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。