首个支持多机械手语义抓取的统一生成框架,让机器人按指令精准抓物。
Multi-GraspLLM: A Multimodal LLM for Multi-Hand Semantic Guided Grasp Generation
- 用大语言模型统一处理不同机械手的抓取序列生成
- 在真实与仿真环境中显著优于现有方法
- 适合研究机器人抓取、多模态生成的开发者
多机械手语义抓取旨在根据自然语言指令为不同机器人手生成可行且语义恰当的抓取姿态。尽管该任务价值高,但由于缺乏带有精细接触描述的多机械手抓取数据集,长期难以突破。本文提出 Multi-GraspSet,首个大规模带自动接触标注的多机械手抓取数据集。基于此,我们构建 Multi-GraspLLM,一种统一的语言引导抓取生成框架,利用大语言模型(LLM)处理变长序列,在单一架构中为多样化机器人手生成抓取姿态。该方法首先将点云特征与文本特征对齐至统一语义空间,再生成抓取二值标记,并通过手感知线性映射转化为各机械手的抓取姿态。实验结果表明,该方法在真实世界与仿真环境中均显著优于现有方法。
原文摘要 · Abstract (English)
Multi-hand semantic grasp generation aims to generate feasible and semantically appropriate grasp poses for different robotic hands based on natural language instructions. Although the task is highly valuable, due to the lack of multihand grasp datasets with fine-grained contact description between robotic hands and objects, it is still a long-standing difficult task. In this paper, we present Multi-GraspSet, the first large-scale multi-hand grasp dataset with automatically contact annotations. Based on Multi-GraspSet, we propose Multi-GraspLLM, a unified language-guided grasp generation framework, which leverages large language models (LLM) to handle variable-length sequences, generating grasp poses for diverse robotic hands in a single unified architecture. Multi-GraspLLM first aligns the encoded point cloud features and text features into a unified semantic space. It then generates grasp bin tokens that are subsequently converted into grasp pose for each robotic hand via hand-aware linear mapping. The experimental results demonstrate that our approach significantly outperforms existing methods in both real-world experiments and simulator. More information can be found on our project page https://multi-graspllm.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。