用大模型规划+可操作性控制,让机器人不抓握也能用工具干活
Non-Prehensile Tool-Object Manipulation by Integrating LLM-Based Planning and Manoeuvrability-Driven Controls
- 用大模型根据语言指令和场景生成可执行的动作序列
- 通过视觉反馈构建工具属性模型,实现狭窄空间内精准操控
- 适合需要灵活使用工具的复杂任务,如工业装配或家庭服务
工具使用曾被视为人类智能的专属能力,但事实上许多动物(如乌鸦)也具备此能力。然而,当前机器人系统在灵巧性方面仍远不及生物体。本文研究利用大语言模型(LLMs)、工具属性信息与物体可操作性,实现非抓握式工具-物体操作。提出一种新方法:基于场景信息与自然语言指令,由大模型生成符号化任务规划,将人类语言指令转化为一系列可行的运动函数。同时设计了一种新的可操作性驱动控制器,通过视觉反馈构建工具属性模型,采用逐步增量策略,在狭小空间内引导机器人完成工具使用与操作。实验验证了该方法在多种操作场景下的有效性。
原文摘要 · Abstract (English)
The ability to wield tools was once considered exclusive to human intelligence, but it's now known that many other animals, like crows, possess this capability. Yet, robotic systems still fall short of matching biological dexterity. In this paper, we investigate the use of Large Language Models (LLMs), tool affordances, and object manoeuvrability for non-prehensile tool-based manipulation tasks. Our novel method leverages LLMs based on scene information and natural language instructions to enable symbolic task planning for tool-object manipulation. This approach allows the system to convert a human language sentence into a sequence of feasible motion functions. We have developed a novel manoeuvrability-driven controller using a new tool affordance model derived from visual feedback. This controller helps guide the robot's tool utilization and manipulation actions, even within confined areas, using a stepping incremental approach. The proposed methodology is evaluated with experiments to prove its effectiveness under various manipulation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。