arXiv:2409.19457cs.RO2024-09ICRA被引 10

用轻量微调让机器人精准听懂指令抓物,适合部署在普通设备上。

A Parameter-Efficient Tuning Framework for Language-guided Object Grounding and Robot Grasping

  • 基于CLIP设计双向视觉语言适配器,实现像素级语义对齐。
  • 引入深度融合分支,提升复杂场景下抓取预测准确率。
  • 仅需少量参数更新,即可支持多任务语言引导抓取。

语言引导的机器人抓取任务要求智能体整合视觉与语言信息,以预测目标驱动的抓取动作。尽管近期基于多模态大模型(MLLM)的方法表现良好,但其高昂的计算和数据需求限制了本地部署与定制化。为此,本文提出一种基于CLIP的轻量级多模态参数高效微调(PET)框架,适用于三种任务:(1)指代表达分割(RES),(2)指代抓取生成(RGS),(3)指代抓取可及性判断(RGA)。该方法创新性地引入双向视觉-语言适配器,实现像素级语言理解对齐,并设计深度融合分支,融合几何线索以辅助抓取预测。实验表明,在RES任务中优于现有基于CLIP的全模型微调或参数高效方法;在RGS与RGA任务中,模型不仅能根据简单语言描述解析物体属性,还展现出对复杂空间推理场景(如工作区存在多个相同物体)的强大理解能力。

原文摘要 · Abstract (English)

The language-guided robot grasping task requires a robot agent to integrate multimodal information from both visual and linguistic inputs to predict actions for target-driven grasping. While recent approaches utilizing Multimodal Large Language Models (MLLMs) have shown promising results, their extensive computation and data demands limit the feasibility of local deployment and customization. To address this, we propose a novel CLIP-based multimodal parameter-efficient tuning (PET) framework designed for three language-guided object grounding and grasping tasks: (1) Referring Expression Segmentation (RES), (2) Referring Grasp Synthesis (RGS), and (3) Referring Grasp Affordance (RGA). Our approach introduces two key innovations: a bi-directional vision-language adapter that aligns multimodal inputs for pixel-level language understanding and a depth fusion branch that incorporates geometric cues to facilitate robot grasping predictions. Experiment results demonstrate superior performance in the RES object grounding task compared with existing CLIP-based full-model tuning or PET approaches. In the RGS and RGA tasks, our model not only effectively interprets object attributes based on simple language descriptions but also shows strong potential for comprehending complex spatial reasoning scenarios, such as multiple identical objects present in the workspace. Project page: https://z.umn.edu/etog-etrg

机器人抓取多模态轻量微调语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。