arXiv:2603.15612cs.CVcs.RO2026-03被引 2

让人体与场景交互重建更真实,直接用于机器人物理仿真。

HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions

论文配图:HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions
图 1 · 摘自论文原文
  • 用物理引擎当监督者,双向优化人体动作和场景结构
  • 在多个交互场景中实现稳定物理仿真,首次达成模拟就可用
  • 适合想把3D重建直接用于机器人控制的研究者

我们提出HSImul3R,一个统一框架,可从非专业拍摄的稀疏视角图像和单目视频中重建逼真且可用于仿真的三维人体-场景交互(HSI)。现有方法存在感知与仿真之间的差距:视觉上合理的重建常违反物理规律,导致物理引擎不稳定,阻碍具身智能应用。为此,我们设计了一种基于物理的双向优化流程,将物理仿真器作为主动监督者,联合优化人体运动与场景几何。正向过程采用面向场景的强化学习,在运动保真度与接触稳定性双重监督下优化人体动作;反向过程提出直接仿真奖励优化,利用重力稳定性和交互成功率反馈来改进场景几何。我们还构建了新基准HSIBench,涵盖多样物体与交互场景。大量实验表明,HSImul3R首次生成稳定、可直接部署的仿真就绪型人体-场景交互重建,可直接应用于真实人形机器人。

原文摘要 · Abstract (English)

We present HSImul3R, a unified framework for simulation-ready 3D reconstruction of human-scene interactions (HSI) from casual captures, including sparse-view images and monocular videos. Existing methods suffer from a perception-simulation gap: visually plausible reconstructions often violate physical constraints, leading to instability in physics engines and failure in embodied AI applications. To bridge this gap, we introduce a physically-grounded bi-directional optimization pipeline that treats the physics simulator as an active supervisor to jointly refine human dynamics and scene geometry. In the forward direction, we employ Scene-targeted Reinforcement Learning to optimize human motion under dual supervision of motion fidelity and contact stability. In the reverse direction, we propose Direct Simulation Reward Optimization, which leverages simulation feedback on gravitational stability and interaction success to refine scene geometry. We further present HSIBench, a new benchmark with diverse objects and interaction scenarios. Extensive experiments demonstrate that HSImul3R produces the first stable, simulation-ready HSI reconstructions and can be directly deployed to real-world humanoid robots.

3D重建物理仿真人机交互机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。