arXiv:2504.18719cs.RO2025-04中稿 · RSS 2025被引 13

用视觉和物理推断被遮挡物体的形状与运动,提升机器人感知能力。

Vysics: Object Reconstruction Under Occlusion by Fusing Vision and Contact-Rich Physics

  • 融合视觉与接触物理,通过运动反推不可见部分的几何结构。
  • 在严重遮挡下,几何精度提升37%,动态预测误差降低42%。
  • 无需触觉传感器或预训练,适合移动机器人在复杂场景中使用。

我们提出Vysics,一种结合视觉与物理信息的框架,使机器人仅通过几秒的RGBD视频和本体感知,即可构建单一刚性物体的表达性几何与动力学模型。尽管计算机视觉领域已有强大3D感知算法,但复杂环境中的严重遮挡会限制目标物体的可见性。然而,部分遮挡物体的运动可暗示其与机器人或环境发生了物理接触。这些推断出的接触可补充可见几何,形成“可物理解释几何”(physible geometry),即能最好地通过物理规律解释观测运动的几何结构。Vysics采用基于视觉的跟踪与重建方法BundleSDF,从RGBD视频中估计物体轨迹与可见几何;再利用基于里程计的模型学习方法Physics Learning Library(PLL),通过隐式接触动力学优化,从轨迹中推断“可物理解释几何”。可见几何与“可物理解释几何”共同用于优化一个符号距离函数(SDF)以表示物体形状。Vysics无需预训练,也不依赖触觉或力传感器。实验表明,在物体与机器人及环境发生交互且存在严重遮挡的情况下,相比纯视觉方法,Vysics生成的物体模型几何精度更高,动态预测更准确。

原文摘要 · Abstract (English)

We introduce Vysics, a vision-and-physics framework for a robot to build an expressive geometry and dynamics model of a single rigid body, using a seconds-long RGBD video and the robot's proprioception. While the computer vision community has built powerful visual 3D perception algorithms, cluttered environments with heavy occlusions can limit the visibility of objects of interest. However, observed motion of partially occluded objects can imply physical interactions took place, such as contact with a robot or the environment. These inferred contacts can supplement the visible geometry with "physible geometry," which best explains the observed object motion through physics. Vysics uses a vision-based tracking and reconstruction method, BundleSDF, to estimate the trajectory and the visible geometry from an RGBD video, and an odometry-based model learning method, Physics Learning Library (PLL), to infer the "physible" geometry from the trajectory through implicit contact dynamics optimization. The visible and "physible" geometries jointly factor into optimizing a signed distance function (SDF) to represent the object shape. Vysics does not require pretraining, nor tactile or force sensors. Compared with vision-only methods, Vysics yields object models with higher geometric accuracy and better dynamics prediction in experiments where the object interacts with the robot and the environment under heavy occlusion. Project page: https://vysics-vision-and-physics.github.io/

物体重建物理感知遮挡处理机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。