arXiv:2602.08058cs.CVcs.AI2026-02中稿 · Robotics: Science …被引 3

让3D场景重建同时满足几何准确与物理合理,避免物体穿插或不稳。

Picasso: Holistic Scene Reconstruction with Physics-Constrained Sampling

  • 基于物理约束的采样方法,结合接触图推理多物体交互
  • 在10个真实场景上实现更符合物理规律的重建结果
  • 适合做仿真规划、机器人抓取等需要物理可信场景的任务

当存在遮挡和测量噪声时,仅几何匹配的场景重建可能在物理上不成立。例如,在模拟器中导入物体位姿和形状估计时,微小误差可能导致物体相互穿透或处于不稳定平衡状态,难以预测接触密集行为的动态演化。为此,本文提出Picasso:一种融合几何、非穿透约束与物理合理性的多物体场景重建框架。该方法采用快速拒绝采样机制,通过推断的物体接触图引导样本生成。同时构建了包含10个接触丰富的真实场景的Picasso数据集,提供真值标注及物理合理性量化指标。在新数据集与YCB-V上的大量实验表明,Picasso显著优于现有方法,重建结果既更贴近真实又符合人类直觉。

原文摘要 · Abstract (English)

In the presence of occlusions and measurement noise, geometrically accurate scene reconstructions -- which fit the sensor data -- can still be physically incorrect. For instance, when estimating the poses and shapes of objects in the scene and importing the resulting estimates into a simulator, small errors might translate to implausible configurations including object interpenetration or unstable equilibrium. This makes it difficult to predict the dynamic behavior of the scene using a digital twin, an important step in simulation-based planning and control of contact-rich behaviors. In this paper, we posit that object pose and shape estimation requires reasoning holistically over the scene (instead of reasoning about each object in isolation), accounting for object interactions and physical plausibility. Towards this goal, our first contribution is Picasso, a physics-constrained reconstruction pipeline that builds multi-object scene reconstructions by considering geometry, non-penetration, and physics. Picasso relies on a fast rejection sampling method that reasons over multi-object interactions, leveraging an inferred object contact graph to guide samples. Second, we propose the Picasso dataset, a collection of 10 contact-rich real-world scenes with ground truth annotations, as well as a metric to quantify physical plausibility, which we open-source as part of our benchmark. Finally, we provide an extensive evaluation of Picasso on our newly introduced dataset and on the YCB-V dataset, and show it largely outperforms the state of the art while providing reconstructions that are both physically plausible and more aligned with human intuition.

场景重建物理约束多物体交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。