arXiv:2605.25751cs.CV2026-05

单图生成高精度可驱动头像,通过分步分裂高斯点提升表情细节。

SplitAvatar: One-shot Head Avatar with Autoregressive Gaussian Splitting

论文配图:SplitAvatar: One-shot Head Avatar with Autoregressive Gaussian Splitting
图 1 · 摘自论文原文
  • 用自回归图网络逐步分裂高斯点,从粗到细重建头部
  • 新密度控制机制避免过度聚集,实现面部区域动态密度调节
  • 适合需要高保真头像的虚拟人、数字孪生应用

3D高斯溅射(3DGS)利用各向异性高斯点实现了高质量场景重建。近期基于3DGS的方法显著提升了人体头像的渲染质量并支持实时性能。然而,现有方法在图像驱动与3DMM驱动生成的高斯点数量上存在数量级差异,导致重建的表情缺乏精细细节。本文提出一种从单张图像重建可动画化头像的新方法。我们设计了图分裂网络,采用自回归架构从粗到细逐步生成高斯点。为解决分裂后图结构不一致问题,引入网格拓扑扩展方法,使GNN连接性与高斯点数量匹配。此外,提出新型密度控制方法,包含门控机制生成软掩码,防止分裂后的过密现象,实现不同面部区域的动态密度控制。为实现平滑快速训练,采用延迟滤波策略,避免训练中重复计算图拓扑。实验表明,自回归结构有效提升了表达表示能力,通过GNN引导的分裂过程合成更精确的面部细节,达到更高重建质量。

原文摘要 · Abstract (English)

3D Gaussian Splatting (3DGS) provides an efficient method for high-quality scene reconstruction using anisotropic Gaussians. Recently, 3DGS-based methods have significantly improved the rendering quality of human avatars while enabling real-time performance. However, existing methods suffer from a magnitude mismatch in the number of Gaussians generated by image-based and 3DMM-based approaches. This discrepancy results in reconstructed expressions that lack fine-grained detail. In this paper, we introduce a novel method for reconstructing an animatable head avatar from a single image. We propose a Graph splitting network to progressively generate Gaussians from coarse to fine using an autoregressive architecture. To address the graph inconsistency caused by split Gaussians, we employ a mesh topology extension method to align the GNN's connectivity with the increased Gaussian count. Furthermore, we introduce a novel density control method that includes a gating mechanism that generates soft masks for Gaussians, preventing over-densification after the splitting operation. This allows for dynamic control over Gaussian density across different facial regions. For smooth and rapid training, we employ a delayed filtering strategy to avoid re-computing the graph topology during training. Experimental results demonstrate that our autoregressive structure effectively improves expression representation ability by progressively splitting Gaussians. This process, enabled by the GNN-guided splitting, synthesizes more precise facial details and achieves higher reconstruction quality.

3D重建高斯溅射头像生成自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。