无需模板的4D人体重建,用视觉先验提升精度与鲁棒性
ShapeGaussian: High-Fidelity 4D Human Reconstruction in Monocular Videos via Vision Priors
- 基于预训练模型学习数据驱动的粗略可变形几何
- 通过神经变形模型捕捉动态细节,实现高保真重建
- 适合需要真实感人体动画的影视、游戏开发者
我们提出ShapeGaussian,一种从普通单目视频中进行高保真4D人体重建的无模板方法。缺乏鲁棒视觉先验的通用重建方法(如4DGS)在无多视角信息时难以捕捉高形变人体运动;而依赖SMPL等模板的方法(如HUGS)虽能生成逼真结果,但对姿态估计错误敏感,常出现不自然伪影。ShapeGaussian有效融合无模板视觉先验,实现高保真且鲁棒的场景重建。方法采用两步流程:首先利用预训练模型估计数据驱动先验,构建粗略可变形几何基础;随后通过神经变形模型细化几何,捕捉细微动态特征。借助2D视觉先验缓解模板方法因姿态估计错误导致的伪影,并使用多参考帧解决无模板方法中2D关键点不可见问题。大量实验表明,ShapeGaussian在重建精度上优于模板方法,在多样人体动作的日常单目视频中均表现更优的视觉质量与鲁棒性。
原文摘要 · Abstract (English)
We introduce ShapeGaussian, a high-fidelity, template-free method for 4D human reconstruction from casual monocular videos. Generic reconstruction methods lacking robust vision priors, such as 4DGS, struggle to capture high-deformation human motion without multi-view cues. While template-based approaches, primarily relying on SMPL, such as HUGS, can produce photorealistic results, they are highly susceptible to errors in human pose estimation, often leading to unrealistic artifacts. In contrast, ShapeGaussian effectively integrates template-free vision priors to achieve both high-fidelity and robust scene reconstructions. Our method follows a two-step pipeline: first, we learn a coarse, deformable geometry using pretrained models that estimate data-driven priors, providing a foundation for reconstruction. Then, we refine this geometry using a neural deformation model to capture fine-grained dynamic details. By leveraging 2D vision priors, we mitigate artifacts from erroneous pose estimation in template-based methods and employ multiple reference frames to resolve the invisibility issue of 2D keypoints in a template-free manner. Extensive experiments demonstrate that ShapeGaussian surpasses template-based methods in reconstruction accuracy, achieving superior visual quality and robustness across diverse human motions in casual monocular videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。