用大模型和多样性算法实现零样本任务感知抓取
Task-Aware Robotic Grasping by evaluating Quality Diversity Solutions through Foundation Models
- 结合语义与几何信息,用LLM指导抓取点选择
- 在65种任务-物体组合中预测准确率达73.6% IoU
- 实测88%用户更偏好任务感知抓取方案
任务感知机器人抓取需融合语义理解与几何推理。本文提出新框架,利用大语言模型(LLMs)与质量多样性(QD)算法,实现零样本任务条件抓取生成。该框架将物体分割为有意义子部件并语义标注,生成结构化表示以提示LLM。通过结合语义与几何表征,使LLM对任务及抓取部位的知识可应用于物理世界。QD生成的抓取档案提供多样化抓取方案,可根据任务选择最优解。在YCB数据集子集上使用Franka Emika机器人评估,通过调查建立任务特定抓取区域的合并真值。方法在65个任务-物体组合中实现73.6%加权交并比(IoU)。小规模端到端验证进一步证实有效性:88%反馈支持任务感知抓取,二项检验显示偏好显著。
原文摘要 · Abstract (English)
Task-aware robotic grasping is a challenging problem that requires the integration of semantic understanding and geometric reasoning. This paper proposes a novel framework that leverages Large Language Models (LLMs) and Quality Diversity (QD) algorithms to enable zero-shot task-conditioned grasp synthesis. The framework segments objects into meaningful subparts and labels each subpart semantically, creating structured representations that can be used to prompt an LLM. By coupling semantic and geometric representations of an object's structure, the LLM's knowledge about tasks and which parts to grasp can be applied in the physical world. The QD-generated grasp archive provides a diverse set of grasps, allowing us to select the most suitable grasp based on the task. We evaluated the proposed method on a subset of the YCB dataset with a Franka Emika robot. A consolidated ground truth for task-specific grasp regions is established through a survey. Our work achieves a weighted intersection over union (IoU) of 73.6% in predicting task-conditioned grasp regions in 65 task-object combinations. An end-to-end validation study on a smaller subset further confirms the effectiveness of our approach, with 88% of responses favoring the task-aware grasp over the control group. A binomial test shows that participants significantly prefer the task-aware grasp.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。