arXiv:2410.01345cs.ROcs.CV2024-10ICRA被引 54

提出新基准与3D+LLM融合策略,提升机器人语言指令泛化能力

Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy

  • 用3D视觉与语言模型结合预测动作,提升已知任务效率
  • 在新场景中表现超越现有方法,尤其在复杂长序列任务上
  • 适合研究机器人泛化、具身智能与多模态决策的团队参考

将语言指令驱动的机器人策略泛化仍面临挑战,主要受限于缺乏合适的仿真评估基准。本文提出GemBench,一个用于评估视觉-语言机器人操作策略泛化能力的新基准。该基准包含七种通用操作原语和四个泛化层级,覆盖新放置位置、刚体与铰接物体以及复杂长时序任务。我们在GemBench上评估了当前先进方法,并引入新模型3D-LOTUS,其利用丰富的3D信息进行语言条件下的动作预测。尽管3D-LOTUS在已见任务中表现出色,但在新任务上表现不佳。为此,我们提出3D-LOTUS++,将3D-LOTUS的运动规划能力与大语言模型(LLM)的任务规划能力及视觉语言模型(VLM)的对象定位精度相结合。3D-LOTUS++在GemBench的新任务上达到最先进水平,为机器人操作的泛化设定新标准。基准数据、代码与训练模型已公开:https://www.di.ens.fr/willow/research/gembench/

原文摘要 · Abstract (English)

Generalizing language-conditioned robotic policies to new tasks remains a significant challenge, hampered by the lack of suitable simulation benchmarks. In this paper, we address this gap by introducing GemBench, a novel benchmark to assess generalization capabilities of vision-language robotic manipulation policies. GemBench incorporates seven general action primitives and four levels of generalization, spanning novel placements, rigid and articulated objects, and complex long-horizon tasks. We evaluate state-of-the-art approaches on GemBench and also introduce a new method. Our approach 3D-LOTUS leverages rich 3D information for action prediction conditioned on language. While 3D-LOTUS excels in both efficiency and performance on seen tasks, it struggles with novel tasks. To address this, we present 3D-LOTUS++, a framework that integrates 3D-LOTUS's motion planning capabilities with the task planning capabilities of LLMs and the object grounding accuracy of VLMs. 3D-LOTUS++ achieves state-of-the-art performance on novel tasks of GemBench, setting a new standard for generalization in robotic manipulation. The benchmark, codes and trained models are available at https://www.di.ens.fr/willow/research/gembench/.

机器人操作视觉语言泛化能力3D规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。