从单目视频重建可模拟的人体与场景,提升交互真实性
CRISP: Contact-Guided Real2Sim from Monocular Video with Planar Scene Primitives
- 用平面几何体拟合点云,通过深度、法向和光流聚类实现清晰场景重建
- 在人体交互遮挡时利用姿态推断被遮部分,运动追踪失败率从55.2%降至6.9%
- 结合强化学习驱动仿人控制器,支持真实视频与AI生成视频的规模化应用
我们提出CRISP,一种从单目视频中恢复可模拟人体运动与场景几何的方法。现有方法依赖数据驱动先验或联合优化,缺乏物理约束,导致几何噪声与伪影,使交互运动策略失效。本工作关键思路是通过简单聚类流程(基于深度、法向与光流)将场景点云拟合为凸形、干净且可用于仿真的平面几何。为重建交互中被遮挡的场景部分,引入人体-场景接触建模(如用人体姿态推断椅子被遮部分)。最终通过强化学习驱动仿人控制器,确保人体与场景重建具有物理合理性。在人体中心视频基准(EMDB、PROX)上,运动追踪失败率由55.2%降至6.9%,强化学习仿真吞吐量提升43%。方法还成功应用于自然拍摄视频、网络视频及Sora生成视频,验证了其大规模生成物理合理人体运动与交互环境的能力,显著推进机器人与AR/VR中的真实世界到仿真应用。
原文摘要 · Abstract (English)
We introduce CRISP, a method that recovers simulatable human motion and scene geometry from monocular video. Prior work on joint human-scene reconstruction relies on data-driven priors and joint optimization with no physics in the loop, or recovers noisy geometry with artifacts that cause motion tracking policies with scene interactions to fail. In contrast, our key insight is to recover convex, clean, and simulation-ready geometry by fitting planar primitives to a point cloud reconstruction of the scene, via a simple clustering pipeline over depth, normals, and flow. To reconstruct scene geometry that might be occluded during interactions, we make use of human-scene contact modeling (e.g., we use human posture to reconstruct the occluded seat of a chair). Finally, we ensure that human and scene reconstructions are physically-plausible by using them to drive a humanoid controller via reinforcement learning. Our approach reduces motion tracking failure rates from 55.2\% to 6.9\% on human-centric video benchmarks (EMDB, PROX), while delivering a 43\% faster RL simulation throughput. We further validate it on in-the-wild videos including casually-captured videos, Internet videos, and even Sora-generated videos. This demonstrates CRISP's ability to generate physically-valid human motion and interaction environments at scale, greatly advancing real-to-sim applications for robotics and AR/VR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。