arXiv:2605.25220cs.CVcs.GR2026-05

仅用单张图像生成高保真3D头像,无需多视角数据或中间渲染。

Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view Generation

论文配图:Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view Generation
图 1 · 摘自论文原文
  • 通过层级双向状态扫描,在3D高斯表示中直接约束多视角一致性。
  • 在真实多视角测试中,纹理和几何一致性均超越现有方法。
  • 适合需要高效构建数字人头像的AR/VR与远程通信场景。

高质量3D高斯头像生成对AR/VR、远程存在和数字人应用至关重要。现有方法依赖多视角数据、3D捕获或中间2D视图合成。本文仅使用随机采样的2D图像,不依赖多视角数据、3D监督或中间视图生成,同时学习条件与非条件3D头像模型。提出MVCHead,一种单次输入的状态空间模型,直接在3D表示中施加多视角一致性(MVC)约束,并在此基础上回归3D高斯。核心是层级状态空间(HiSS)块,从粗到细逐步优化高斯,捕捉长程依赖;在每个HiSS块中,将Mamba的单向扫描改进为层级双向状态扫描(HiBiSS),使递归沿多视角不一致最强的轴对齐。最后设计SE(3)多视角判别器,判断自渲染集合是否源自单一3D配置,通过奖励跨视角像素对齐来评估,无需真实多视角配对。MVCHead在感知质量上达到当前最优,纹理与几何一致性优于先前方法,形状一致性保持相当。为验证可扩展性,发布FaceGS-10K,首个大规模现成可用的3D高斯头像资产数据集,用于训练与评估3D头像模型。

原文摘要 · Abstract (English)

High-fidelity 3D Gaussian head avatar generation is critical for applications such as AR/VR, telepresence, and digital humans. Existing methods depend on multi-view datasets, 3D captures, or intermediate 2D view synthesis. In contrast, we learn both conditional and unconditional 3D head models from randomly sampled 2D images alone, without using multi-view data, 3D supervision, or intermediate view generation. We introduce MVCHead, a single-shot state space model that enforces multi-view consistency (MVC) directly in the 3D representation while regressing 3D Gaussians under these constraints. At its core, we propose a Hierarchical State Space (HiSS) block that progressively refines Gaussians from coarse to fine, while capturing long-range dependencies. Within each HiSS block, we modify Mamba's standard unidirectional scan with the proposed Hierarchical Bi-directional State Scan (HiBiSS) that aligns recurrence with the axes along which multi-view inconsistencies are strongest. Finally, we design an SE(3) Multi-view Critic that judges whether a set of self-renders arises from a single underlying 3D configuration, rewarding cross-view pixel alignment without observing real multi-view pairs. MVCHead achieves state-of-the-art perceptual quality, surpasses prior methods in both texture and geometric consistency, and maintains comparable shape consistency. To demonstrate scalability, we release FaceGS-10K, the first large-scale dataset of ready-to-use 3D Gaussian head assets for training and evaluation of 3D head models. Project Page and code: https://humansensinglab.github.io/MVCHead/

3D生成高斯漫游单图建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。