从单目视频中零样本重建物体在场景中的操作过程。
Zero-shot Reconstruction of In-Scene Object Manipulation from Video
- 用基础模型初始化物体、场景和手部姿态
- 两阶段优化恢复从抓取到交互的完整动作序列
- 无需标注数据,适合真实场景应用
我们构建了首个从单目RGB视频中重建场景内物体操作的系统。该问题因场景重建病态、手物深度模糊以及需物理合理交互而极具挑战。现有方法基于手部坐标系,忽略场景信息,影响度量精度与实用性。我们的方法首先利用数据驱动的基础模型初始化物体网格与姿态、场景点云及手部姿态;随后通过两阶段优化,恢复从抓取到交互的完整手物运动,且与输入视频中的场景信息保持一致。
原文摘要 · Abstract (English)
We build the first system to address the problem of reconstructing in-scene object manipulation from a monocular RGB video. It is challenging due to ill-posed scene reconstruction, ambiguous hand-object depth, and the need for physically plausible interactions. Existing methods operate in hand centric coordinates and ignore the scene, hindering metric accuracy and practical use. In our method, we first use data-driven foundation models to initialize the core components, including the object mesh and poses, the scene point cloud, and the hand poses. We then apply a two-stage optimization that recovers a complete hand-object motion from grasping to interaction, which remains consistent with the scene information observed in the input video.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。