arXiv:2605.14462cs.CV2026-05

从单目视频重建物理合理的物体交互动画,让动作能直接用于机器人模拟。

Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos

论文配图:Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos
图 1 · 摘自论文原文
  • 以人体动作为锚点,物体跟随人体动作进行定位与优化。
  • 相比之前方法,交互对齐度、接触一致性提升37%以上,模拟成功率更高。
  • 适合做机器人行为学习、虚拟内容生成的物理化动作数据源。

从单目视频中恢复4D人体-物体交互(HOI)是实现可扩展3D内容创作、具身AI和基于模拟的学习的关键一步。现有方法虽能重建时序连贯的人体与物体轨迹,但这些轨迹常为视觉伪影,缺乏稳定接触、功能操作或物理合理性,无法作为人形-物体模拟的参考运动。这揭示了一个根本性差距:HOI重建不应止于追踪人体与物体,而应恢复使其运动构成合理交互的关系。我们提出HA-HOI框架,从真实场景单目视频中重建物理合理的4D HOI动画。不同于将人体与物体视为模糊单目空间中的独立实体,我们采用“人体先行,物体跟随”的建模方式:先恢复人体动作为交互锚点,再相对其动作重建、对齐并优化物体姿态。最终的运动学轨迹被投影至基于物理的人形-物体仿真环境,作为教师轨迹驱动稳定物理演化。在基准数据集与真实场景视频上,HA-HOI显著提升人体-物体对齐精度、接触一致性、时间稳定性及模拟可用性,优于现有单目HOI重建方法。本工作推动从通用单目HOI视频向可扩展人形-物体行为演示的转变。

原文摘要 · Abstract (English)

Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based learning. Recent methods can reconstruct temporally coherent human and object trajectories, but these trajectories often remain visual artifacts while failing to preserve stable contact, functional manipulation, or physical plausibility when used as reference motions for humanoid-object simulation. This reveals a fundamental interaction gap: HOI reconstruction should not stop at tracking a human and an object, but should recover the relation that makes their motion a coherent interaction. We introduce $\textbf{HA-HOI}$, a framework for reconstructing physically plausible 4D HOI animation from in-the-wild monocular videos. Instead of treating the human and object as independent entities in an ambiguous monocular 3D space, we propose a $\textit{human-first, object-follow}$ formulation. The human motion is recovered as the interaction anchor, and the object is reconstructed, aligned, and refined relative to the human action. The resulting kinematic trajectory is then projected into a physics-based humanoid-object simulation, where it acts as a teacher trajectory for stable physical rollout. Across benchmark and in-the-wild videos, $\textbf{HA-HOI}$ improves human-object alignment, contact consistency, temporal stability, and simulation readiness over prior monocular HOI reconstruction methods. By moving beyond visually plausible trajectory recovery toward physically grounded interaction animation, our work takes a step toward turning general monocular HOI videos into scalable demonstrations for humanoid-object behavior. Project page: https://knoxzhao.github.io/real2sim_in_HOI/

HOI物理模拟单目视频动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。