arXiv:2606.05359cs.CV2026-06

用物理仿真优化单目视频中人物与物体交互的合理性。

Recovering Physically Plausible Human-Object Interactions from Monocular Videos

论文配图:Recovering Physically Plausible Human-Object Interactions from Monocular Videos
图 1 · 摘自论文原文
  • 先用运动学估计初始动作,再通过强化学习优化。
  • 在物理模拟器中训练策略,显著减少穿模和漂浮现象。
  • 适合需要真实物理交互的动画生成与动作分析场景。

本文提出 RePHO,一种从单目视频中重建物理合理的人物-物体交互(HOI)的方法。现有基于运动学的方法虽能生成视觉上合理的动作,但常导致穿模、物体漂浮等物理不合理现象。为此,我们设计了一种物理引导的重构框架:先获取运动学估计,再通过强化学习(RL)训练策略,使其在物理模拟器中重现交互过程。由于运动学估计通常存在噪声,直接进行RL训练易失败。因此,我们提出一种自适应采样策略,结合双自更新机制,识别出最具信息量且可靠的重建帧。该方法逐步提升重建质量,生成物理一致的HOI序列。我们在两个标准的HOI基准上验证了该方法,相比现有最优方法,在物理合理性指标上取得明显提升。

原文摘要 · Abstract (English)

In this paper, we propose RePHO, a method to reconstruct physically plausible human-object interactions (HOI) from monocular videos. While existing kinematic-based approaches produce visually plausible motion, they often result in physically implausible artifacts such as interpenetration and object floating. To overcome these issues, we introduce a physics-guided reconstruction framework. We begin with a kinematic estimate and then refine it by training a policy with reinforcement learning (RL). This policy is optimized to reproduce the interaction in a physics simulator. Because kinematic estimates are typically noisy, naive RL training can fail. Therefore, we propose an adaptive sampling strategy with a dual self-updating mechanism that can identify the frames with the most informative and reliable kinematic reconstruction. Our process progressively improves reconstruction quality and yields physically consistent HOI sequences. We demonstrate our approach on two standard HOI benchmarks and achieve clear improvements in physical plausibility metrics over state-of-the-art methods. Project Page: https://dingbang777.github.io/RePHO/

人物交互物理仿真单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。