用网格引导的2D高斯点云+大模型先验,实现单目视频高保真人物重建
FMGS-Avatar: Mesh-Guided 2D Gaussian Splatting with Foundation Model Priors for 3D Monocular Avatar Reconstruction
- 将2D高斯点绑定到模板网格面片,约束位置旋转,提升表面贴合度
- 融合大模型先验知识,显著改善单视角下的几何与外观还原效果
- 适合需要高质量可驱动虚拟人像的影视、游戏与元宇宙场景
从单目视频重建高保真可驱动的人类虚拟形象仍具挑战,源于单视角观测中几何信息不足。尽管近期3D高斯溅射方法表现良好,但其自由形态的3D高斯原语难以保持表面细节。为此,我们提出新方法FMGS-Avatar,集成两项关键创新:首先,引入网格引导的2D高斯溅射,将2D高斯原语直接附着于模板网格面片,通过约束位置、旋转和运动,实现更优的表面对齐与几何细节保留;其次,利用在大规模数据集上训练的基础模型(如Sapiens)补充单目视频有限的视觉线索。然而,在从基础模型中蒸馏多模态先验时,不同模态表现出不同的参数敏感性,导致优化目标冲突。我们通过选择性梯度隔离的协同训练策略,使各损失组件仅优化相关参数,避免干扰。结合增强表示与协调信息蒸馏,本方法显著提升单目人体虚拟形象重建质量。实验表明,相比现有方法,本方法在几何精度与外观保真度上均有明显提升,并能提供丰富的语义信息。此外,共享规范空间中的先验知识天然支持新视角与新姿态下时空一致的渲染。
原文摘要 · Abstract (English)
Reconstructing high-fidelity animatable human avatars from monocular videos remains challenging due to insufficient geometric information in single-view observations. While recent 3D Gaussian Splatting methods have shown promise, they struggle with surface detail preservation due to the free-form nature of 3D Gaussian primitives. To address both the representation limitations and information scarcity, we propose a novel method, \textbf{FMGS-Avatar}, that integrates two key innovations. First, we introduce Mesh-Guided 2D Gaussian Splatting, where 2D Gaussian primitives are attached directly to template mesh faces with constrained position, rotation, and movement, enabling superior surface alignment and geometric detail preservation. Second, we leverage foundation models trained on large-scale datasets, such as Sapiens, to complement the limited visual cues from monocular videos. However, when distilling multi-modal prior knowledge from foundation models, conflicting optimization objectives can emerge as different modalities exhibit distinct parameter sensitivities. We address this through a coordinated training strategy with selective gradient isolation, enabling each loss component to optimize its relevant parameters without interference. Through this combination of enhanced representation and coordinated information distillation, our approach significantly advances 3D monocular human avatar reconstruction. Experimental evaluation demonstrates superior reconstruction quality compared to existing methods, with notable gains in geometric accuracy and appearance fidelity while providing rich semantic information. Additionally, the distilled prior knowledge within a shared canonical space naturally enables spatially and temporally consistent rendering under novel views and poses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。