arXiv:2506.11261cs.ROcs.AI2025-06被引 2

让机器人看懂语言指令并精准操作新物体,跨场景泛化能力强

Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation

  • 用多视角图像和历史动作生成带语义标注的下一步计划
  • 在GemBench上四种泛化场景下均超越当前最佳方法
  • 适合研究通用机器人操控与视觉语言导航的学者

机器人操控在面对未见物体、环境及多样化语言指令时面临泛化难题。现有基于大语言模型的方法虽具潜力,但难以在视觉环境中生成准确的具身规划。尽管已有工作尝试对语言模型进行视觉指令微调,但通常受限于单视角输入,难以实现精确物体定位。本文提出Gondola,一种基于大语言模型的新型具身视觉-语言规划模型,可接收多视角图像与历史计划,输出包含文本描述与目标物体/位置分割掩码的交错式动作规划。为支持训练,我们利用RLBench模拟器构建三类数据集:机器人具身规划、多视角指代表达与伪长时任务数据集。Gondola在GemBench数据集的四个泛化层级(新放置、刚性物体、铰接物体、长时任务)上均优于现有最先进方法。

原文摘要 · Abstract (English)

Robotic manipulation faces a significant challenge in generalizing across unseen objects, environments and tasks specified by diverse language instructions. To improve generalization capabilities, recent research has incorporated large language models (LLMs) for planning and action execution. While promising, these methods often fall short in generating grounded plans in visual environments. Although efforts have been made to perform visual instructional tuning on LLMs for robotic manipulation, existing methods are typically constrained by single-view image input and struggle with precise object grounding. In this work, we introduce Gondola, a novel grounded vision-language planning model based on LLMs for generalizable robotic manipulation. Gondola takes multi-view images and history plans to produce the next action plan with interleaved texts and segmentation masks of target objects and locations. To support the training of Gondola, we construct three types of datasets using the RLBench simulator, namely robot grounded planning, multi-view referring expression and pseudo long-horizon task datasets. Gondola outperforms the state-of-the-art LLM-based method across all four generalization levels of the GemBench dataset, including novel placements, rigid objects, articulated objects and long-horizon tasks.

机器人操控视觉语言大模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。