arXiv:2606.15659cs.CV2026-06

用统一表示法实现高质量4D人脸重建,零样本跨域表现领先。

SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction

论文配图:SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction
图 1 · 摘自论文原文
  • 基于FLAME网格绑定的高斯表示,分阶段融合单目与多视角信息。
  • 跨域零样本测试下比GAGAvatar高1.5 dB PSNR,单目基准领先所有现有方法。
  • 仅需10K次迭代即可完成精修,速度比主流方案快60倍,适合实时应用。

从一张或少数几张源肖像生成高质量4D头部虚拟人是远程通信、增强现实/虚拟现实及数字人交互的核心。3D高斯点云(3DGS)已成为主流表示方式,其包含通用前馈预测器与个体微调器两类范式。但现有前馈模型依赖单一数据集且固定输入数量,存在领域偏差;而个体微调器需30万至60万次迭代,依赖自适应加密会破坏原有高斯布局,无法端到端共享表示。为此,本文提出SpatialAvatar-0,采用共享的FLAME网格绑定高斯表示:一个无参数的K源均值池化前馈生成器,结合单目-时序到多视图-空间的两阶段调度策略,防止身份先验坍塌;并引入10,000次迭代的布局保持微调循环,冻结FLAME绑定与高斯数量,以三组件抗突增正则替代加密。在VFHQ/HDTF跨域零样本测试中,优于在域内领先的GAGAvatar达+1.5 dB PSNR,且从未训练于任一测试域;在SplattingAvatar单目基准上,各项指标全面领先,超越30万次迭代的GeoAvatar达+1.3 dB PSNR,且微调耗时最多缩短60倍。

原文摘要 · Abstract (English)

High-quality 4D head avatars from one or a few source portraits are central to telepresence, AR/VR, and digital-human interaction. 3D Gaussian Splatting (3DGS) has emerged as the dominant representation, with two complementary regimes (generalizable feed-forward predictors and per-subject refiners) maturing in parallel. However, existing feed-forward predictors are trained on a single dataset family with a hard-coded source count, inheriting the corresponding domain bias. Per-subject refiners require 300K--600K iterations and rely on adaptive densification that destroys upstream Gaussian layouts, preventing the two regimes from sharing a representation end-to-end. To bridge both regimes we propose SpatialAvatar-0 on a shared FLAME-mesh-bound Gaussian representation: a feed-forward generator with a parameter-free K-source mean-pool and a monocular-temporal to multi-view-spatial two-phase schedule that anchors against identity-prior collapse onto the smaller multi-view set. We further introduce a 10K-iter layout-preserving per-subject refinement loop that freezes the FLAME-binding and Gaussian count and replaces densification with a three-component anti-spike regularization. On VFHQ/HDTF cross-domain zero-shot we surpass the in-domain leader GAGAvatar by +1.5 dB PSNR despite never training on either test domain, and on the SplattingAvatar monocular benchmark we lead every reported metric, surpassing the 300K-iter GeoAvatar by +1.3 dB PSNR at up to 60x shorter per-subject schedule than common SOTA baselines. Website: https://spatialwalk.github.io/SpatialAvatar-0.

4D重建高斯点云虚拟人实时生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。