arXiv:2601.13837cs.CV2026-01中稿 · ICLR被引 6

仅用几张图生成高保真3D人脸,实时驱动动画

FastGHA: Generalized Few-Shot 3D Gaussian Head Avatars with Real-Time Animation

  • 直接从输入图像学习像素级高斯表示,融合多视角特征
  • 生成效果优于现有方法,推理速度更快,支持实时动画
  • 适合快速构建可动3D人脸,尤其适用于新用户

尽管基于3D高斯的人脸建模取得进展,高效生成高质量头像仍是挑战。现有方法通常依赖大量多视角拍摄或单目视频配合身份专属优化,限制了对未见主体的可扩展性和易用性。为此,我们提出FastGHA,一种前馈方法,仅需少量输入图像即可生成高质量高斯头像,并支持实时动画。该方法直接从输入图像学习每像素高斯表示,利用基于Transformer的编码器融合DINOv3和Stable Diffusion VAE的图像特征以聚合多视角信息。为实现实时动画,我们在显式高斯表示中引入每高斯特征,并设计轻量级MLP动态网络,从表情代码预测3D高斯形变。此外,为提升头部几何平滑性,我们采用预训练大模型生成的点云图作为几何监督。实验表明,该方法在渲染质量与推理效率上显著优于现有方法,同时支持实时动态头像动画。

原文摘要 · Abstract (English)

Despite recent progress in 3D Gaussian-based head avatar modeling, efficiently generating high fidelity avatars remains a challenge. Current methods typically rely on extensive multi-view capture setups or monocular videos with per-identity optimization during inference, limiting their scalability and ease of use on unseen subjects. To overcome these efficiency drawbacks, we propose FastGHA, a feed-forward method to generate high-quality Gaussian head avatars from only a few input images while supporting real-time animation. Our approach directly learns a per-pixel Gaussian representation from the input images, and aggregates multi-view information using a transformer-based encoder that fuses image features from both DINOv3 and Stable Diffusion VAE. For real-time animation, we extend the explicit Gaussian representations with per-Gaussian features and introduce a lightweight MLP-based dynamic network to predict 3D Gaussian deformations from expression codes. Furthermore, to enhance geometric smoothness of the 3D head, we employ point maps from a pre-trained large reconstruction model as geometry supervision. Experiments show that our approach significantly outperforms existing methods in both rendering quality and inference efficiency, while supporting real-time dynamic avatar animation.

3D头像实时动画少样本高斯建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。