arXiv:2502.13964cs.ROcs.AI2025-02被引 2

用视觉模型实现机器人精准抓取小物件,无需预先训练。

Precise Mobile Manipulation of Small Everyday Objects

  • 利用视觉大模型生成3D目标,闭环控制机械臂完成精细操作。
  • 在10个新环境中对72个物体测试,零样本成功率71%,比基线高50%。
  • 通过图像修复遮挡的末端执行器,提升目标定位精度,适合真实场景应用。

许多日常移动操作任务需要精确处理小物体,如旋转门把手或按开关。本文提出视觉模型引导的伺服控制(SVM),一种闭环框架,使移动机械臂能在新环境中完成此类高精度任务。SVM利用先进的视觉基础模型生成用于视觉伺服的3D目标,实现多样任务。由于末端执行器遮挡导致目标定位失败,SVM通过视觉模型对遮挡区域进行补全,显著提升定位效果。实验表明,结合开放词汇检测器可自动识别语义目标(如门把手),结合点追踪方法可准确响应用户点击的位置。我们在6栋建筑中的10个新环境里进行了大规模评估,共涉及72种不同物体。SVM在真实世界中对未见过的物体实现71%的零样本成功率,比开环控制方法高42个百分点,也优于在1000多个示范数据上训练的模仿学习基线,成功率达50个百分点。

原文摘要 · Abstract (English)

Many everyday mobile manipulation tasks require precise interaction with small objects, such as grasping a knob to open a cabinet or pressing a light switch. In this paper, we develop Servoing with Vision Models (SVM), a closed-loop framework that enables a mobile manipulator to tackle such precise tasks involving the manipulation of small objects. SVM uses state-of-the-art vision foundation models to generate 3D targets for visual servoing to enable diverse tasks in novel environments. Naively doing so fails because of occlusion by the end-effector. SVM mitigates this using vision models that out-paint the end-effector, thereby significantly enhancing target localization. We demonstrate that aided by out-painting methods, open-vocabulary object detectors can serve as a drop-in module for SVM to seek semantic targets (e.g. knobs) and point tracking methods can help SVM reliably pursue interaction sites indicated by user clicks. We conduct a large-scale evaluation spanning experiments in 10 novel environments across 6 buildings including 72 different object instances. SVM obtains a 71% zero-shot success rate on manipulating unseen objects in novel environments in the real world, outperforming an open-loop control method by an absolute 42% and an imitation learning baseline trained on 1000+ demonstrations also by an absolute success rate of 50%.

移动操作视觉伺服小物体零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。