仅用少量演示即可让机器人理解语言指令并完成新环境中的物品摆放任务。
Improving Generalization of Language-Conditioned Robot Manipulation
- 分两阶段处理:先定位目标物,再确定放置区域。
- 通过实例级语义融合,精准匹配语言描述与图像中的物体。
- 在真实机器人上实现零样本泛化,显著减少数据依赖。
机器人操控通常依赖视觉输入。近年来,视觉语言模型(VLMs)的发展使得自然语言指令可用来指导视觉输入并控制机器人在更广泛环境中运行。然而,现有方法需大量数据微调VLM以适应未见过的环境。本文提出一种框架,仅需少量示范即可学习物体排列任务。我们设计了两阶段流程:第一阶段定位目标物体,第二阶段确定放置区域。引入实例级语义融合模块,将图像实例切片与文本嵌入对齐,使模型能根据自然语言指令识别目标物体。我们在仿真和真实机器人环境中验证了该方法。仅用少量示范微调后,该方法显著提升泛化能力,并在真实机器人场景中实现零样本操作。
原文摘要 · Abstract (English)
The control of robots for manipulation tasks generally relies on visual input. Recent advances in vision-language models (VLMs) enable the use of natural language instructions to condition visual input and control robots in a wider range of environments. However, existing methods require a large amount of data to fine-tune VLMs for operating in unseen environments. In this paper, we present a framework that learns object-arrangement tasks from just a few demonstrations. We propose a two-stage framework that divides object-arrangement tasks into a target localization stage, for picking the object, and a region determination stage for placing the object. We present an instance-level semantic fusion module that aligns the instance-level image crops with the text embedding, enabling the model to identify the target objects defined by the natural language instructions. We validate our method on both simulation and real-world robotic environments. Our method, fine-tuned with a few demonstrations, improves generalization capability and demonstrates zero-shot ability in real-robot manipulation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。