arXiv:2602.24161cs.CV2026-02

用几何感知扩散模型重建高保真可动画4D人脸头像

GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar Reconstruction

  • 联合生成图像与法向图,学习强几何先验
  • 在视觉质量、表情还原和跨身份泛化上超越现有方法
  • 支持实时渲染,适合虚拟人与元宇宙应用

从单张正面照重建逼真且可动画的4D人脸头像仍是计算机视觉中的基本挑战。尽管扩散模型在头像重建的图像与视频生成中取得显著进展,现有方法主要依赖2D先验,难以保持一致的3D几何结构。我们提出一种新型框架,利用几何感知扩散学习强几何先验,以实现高保真头像重建。该方法联合生成肖像图像与对应的表面法向图,同时通过无姿态的表情编码器捕捉隐式表情表征。合成的图像与表情潜在表示被整合进基于3D高斯的头像中,实现高保真渲染与精确几何重建。大量实验表明,本方法在视觉质量、表情保真度及跨身份泛化能力上显著优于当前最优方法,同时支持实时渲染。

原文摘要 · Abstract (English)

Reconstructing photorealistic and animatable 4D head avatars from a single portrait image remains a fundamental challenge in computer vision. While diffusion models have enabled remarkable progress in image and video generation for avatar reconstruction, existing methods primarily rely on 2D priors and struggle to achieve consistent 3D geometry. We propose a novel framework that leverages geometry-aware diffusion to learn strong geometry priors for high-fidelity head avatar reconstruction. Our approach jointly synthesizes portrait images and corresponding surface normals, while a pose-free expression encoder captures implicit expression representations. Both synthesized images and expression latents are incorporated into 3D Gaussian-based avatars, enabling photorealistic rendering with accurate geometry. Extensive experiments demonstrate that our method substantially outperforms state-of-the-art approaches in visual quality, expression fidelity, and cross-identity generalization, while supporting real-time rendering.

4D重建扩散模型3D高斯虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。