arXiv:2605.02784cs.CV2026-05

通过闭环优化提升视频中人体姿态与外观重建精度

HumanSplatHMR: Closing the Loop Between Human Mesh Recovery and Gaussian Splatting Avatar

论文配图:HumanSplatHMR: Closing the Loop Between Human Mesh Recovery and Gaussian Splatting Avatar
图 1 · 摘自论文原文
  • 联合优化人体姿态与高保真姿态,实现端到端的3D人体重建
  • 在公开数据集上显著提升姿态估计准确率,新视角渲染效果更真实
  • 适合野外场景下无需运动捕捉的虚拟人像生成应用

从视频中精确恢复人体姿态与外观是场景重建的关键,广泛应用于动作捕捉、虚拟现实和数字孪生。现有方法在三维人体几何重建上表现不佳:基于ViT的方法不稳定且易过拟合2D视角;而基于NeRF或高斯溅射的化身则将姿态与外观分开处理,限制了新姿态下的渲染泛化能力。为此,本文提出HumanSplatHMR,一种联合优化框架,在同时学习高保真化身的同时精修3D人体姿态。核心思想是建立几何姿态估计与可微渲染之间的闭环反馈。不同于依赖运动捕捉系统或离线优化的姿态输入,本方法仅使用先进人体姿态估计算器输出的网格估计,更贴近真实场景。通过在可微渲染器中反向传播光度、分割和深度损失至姿态参数与全局位置,实现姿态随时间迭代优化,提升精度与对齐性,并改善新视角渲染效果。实验表明,该方法在姿态恢复与化身生成方面均优于忽略图像级优化的基线模型和解耦姿态与重建的基准方法。

原文摘要 · Abstract (English)

Accurately recovering human pose and appearance from video is an essential component of scene reconstruction, with applications to motion capture, motion prediction, virtual reality, and digital twinning. Despite significant interest in building realistic human avatars from video, this paper demonstrates that existing methods do not accurately recover the 3D geometry of humans. ViT-based approaches are not consistently reliable and can overfit to 2D views, while NeRF- and Gaussian Splatting-based avatars treat pose and appearance separately, limiting rendering generalization to new poses. To resolve these shortcomings, this paper proposes HumanSplatHMR, a joint optimization framework that refines 3D human poses while simultaneously learning a high-fidelity avatar for novel-view and novel-pose synthesis. Our key insight is to close the loop between geometric pose estimation and differentiable rendering. Unlike prior human avatar methods that rely on accurate human pose obtained through motion capture systems or offline refinement, which are impractical in in-the-wild scenarios, our approach uses only human mesh estimates from a state-of-the-art human pose estimator to better reflect real-world conditions. Therefore, instead of using the human pose only as a deformation prior, HumanSplatHMR backpropagates photometric, segmentation, and depth losses through a differentiable renderer to the pose parameters and global position. This coupling refines the global 3D pose over time, improving accuracy and alignment while producing better renderings from novel views. Experiments show consistent improvements over pose recovery baselines that omit image-level refinement and avatar baselines that decouple pose estimation from avatar reconstruction.

人体重建高斯溅射姿态优化虚拟人像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。