arXiv:2606.12910cs.ROcs.AI2026-06

用语言指令控制机器人抓取,无需训练就能理解'上层货架'等抽象指令。

Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning

论文配图:Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning
图 1 · 摘自论文原文
  • 用视觉语言模型将语言转为带边框的符号化目标,实现物理世界对齐。
  • 在三种难度下90次真实机器人测试中成功率73.3%,无需任务微调。
  • 适合需要快速适应新指令的家用或工业场景机器人应用。

为使机器人有效融入家庭或工业环境,机器需能实时响应自然语言指令。尽管视觉语言模型(VLM)已实现机器人任务与运动规划(TAMP)的零样本泛化,但现有顶尖方法通常计算开销大或需数千次示范训练。本文提出GRASP(Grounded Reasoning and Symbolic Planning)框架,旨在实现开放词汇的桌面操作。该方法利用预训练VLM将自然语言查询转化为神经符号目标状态,并通过边界框检测流程在物理世界中实现语义锚定。相比依赖固定颜色列表或硬编码坐标的方案,GRASP可解析“上层货架”等抽象空间概念并执行任务而无需额外微调。在三个难度等级的90次真实机器人试验中,整体成功率达73.3%,且无需任务特定训练。

原文摘要 · Abstract (English)

For robotics to be effectively integrated into household or industrial environments, machines must adapt to natural-language prompts in real time. Although Vision-Language Models (VLMs) have enabled zero-shot generalization in robot task and motion planning (TAMP), current state-of-the-art approaches often remain computationally "heavyweight" or require extensive training on thousands of demonstrations. We present GRASP (Grounded Reasoning and Symbolic Planning), a framework designed as a step toward open-vocabulary tabletop manipulation. Our approach leverages a pretrained VLM to translate natural-language queries into neuro-symbolic goal states, grounded in the physical world via a bounding-box detection pipeline. Unlike methods that rely on fixed color lists or hard-coded coordinates, GRASP enables robots to interpret abstract spatial concepts such as "top shelf" and execute tasks without additional fine-tuning. We achieve 73.3% overall success across 90 real-robot trials at three difficulty levels, requiring no task-specific training.

机器人抓取语言理解神经符号零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。