让机器人听懂自然语言,完成复杂抓取任务
RoboDexVLM: Visual Language Model-Enabled Task Planning and Motion Control for Dexterous Robot Manipulation
- 用视觉语言模型理解自然语言指令,规划长序列操作
- 零样本适配多种物体,实现多样抓取动作
- 适合需要灵活操作的工业与服务场景
本文提出RoboDexVLM框架,面向配备灵巧手的协作机器人,实现基于自然语言指令的长时序任务规划与抓取感知。现有方法多局限于简化任务,忽视复杂环境下的多样化抓取需求。本框架通过视觉语言模型(VLM)构建具备任务级容错能力的任务规划器,可执行开放词汇命令;同时设计语言引导的灵巧抓取感知算法,结合机器人运动学与形式化方法,支持零样本条件下对多种形状物体的灵巧操作。实验验证了该框架在长时序任务和复杂环境中的有效性、适应性与鲁棒性,展现出其在开放词汇灵巧操纵中的潜力。项目开源地址:https://henryhcliu.github.io/robodexvlm。
原文摘要 · Abstract (English)
This paper introduces RoboDexVLM, an innovative framework for robot task planning and grasp detection tailored for a collaborative manipulator equipped with a dexterous hand. Previous methods focus on simplified and limited manipulation tasks, which often neglect the complexities associated with grasping a diverse array of objects in a long-horizon manner. In contrast, our proposed framework utilizes a dexterous hand capable of grasping objects of varying shapes and sizes while executing tasks based on natural language commands. The proposed approach has the following core components: First, a robust task planner with a task-level recovery mechanism that leverages vision-language models (VLMs) is designed, which enables the system to interpret and execute open-vocabulary commands for long sequence tasks. Second, a language-guided dexterous grasp perception algorithm is presented based on robot kinematics and formal methods, tailored for zero-shot dexterous manipulation with diverse objects and commands. Comprehensive experimental results validate the effectiveness, adaptability, and robustness of RoboDexVLM in handling long-horizon scenarios and performing dexterous grasping. These results highlight the framework's ability to operate in complex environments, showcasing its potential for open-vocabulary dexterous manipulation. Our open-source project page can be found at https://henryhcliu.github.io/robodexvlm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。