arXiv:2501.02158cs.CV2025-01被引 31

从单目视频中联合优化人体与场景的4D重建,提升自然场景下动作还原精度。

Joint Optimization for 4D Human-Scene Reconstruction in the Wild

  • 基于人体-场景接触约束,联合优化场景、相机位姿与人体运动
  • 在自然场景视频上实现更准确的人体动作与稠密场景重建
  • 新模型JOSH3R仅用伪标签训练即超越现有无优化方法

重建人体运动及其周围环境对于理解人-场景交互和预测行为至关重要。尽管已有方法在受控环境下取得进展,但难以从网络视频中还原自然多样的人体动作与场景上下文。本文提出JOSH,一种基于优化的野外单目视频4D人体-场景重建方法。JOSH结合密集场景重建与人体网格恢复技术进行初始化,并利用人体-场景接触约束,联合优化场景几何、相机位姿与人体运动。实验表明,通过联合优化,JOSH在全局人体运动估计与稠密场景重建方面均表现更优。我们进一步设计了高效模型JOSH3R,直接使用网络视频中的伪标签进行训练。JOSH3R仅凭JOSH生成的预测标签即可超越其他无优化方法,证明其高精度与强泛化能力。

原文摘要 · Abstract (English)

Reconstructing human motion and its surrounding environment is crucial for understanding human-scene interaction and predicting human movements in the scene. While much progress has been made in capturing human-scene interaction in constrained environments, those prior methods can hardly reconstruct the natural and diverse human motion and scene context from web videos. In this work, we propose JOSH, a novel optimization-based method for 4D human-scene reconstruction in the wild from monocular videos. JOSH uses techniques in both dense scene reconstruction and human mesh recovery as initialization, and then it leverages the human-scene contact constraints to jointly optimize the scene, the camera poses, and the human motion. Experiment results show JOSH achieves better results on both global human motion estimation and dense scene reconstruction by joint optimization of scene geometry and human motion. We further design a more efficient model, JOSH3R, and directly train it with pseudo-labels from web videos. JOSH3R outperforms other optimization-free methods by only training with labels predicted from JOSH, further demonstrating its accuracy and generalization ability.

4D重建人体动作单目视频联合优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。