arXiv:2409.14066cs.ROcs.AI2024-09ICRA被引 19

用50个例子让机器人学会操作新物体,无需真实机械臂数据

KALIE: Fine-Tuning Vision-Language Models for Open-World Manipulation without Robot Data

  • 用视觉语言模型预测关键点可达性,实现自然语言控制
  • 仅需50个示例即可让机器人在新任务中稳定操作未知物体
  • 自动生成高质量训练数据,避免依赖真实机器人采集

构建通用机器人系统需要让机器人具备在开放世界中处理新物体的能力。受大型预训练模型进展的启发,我们提出基于想象环境的关键点可操作性学习(KALIE),以可扩展方式适配预训练视觉语言模型(VLM)用于机器人控制。KALIE不直接生成运动指令,而是根据自然语言指令和场景视觉观察,预测基于点的可操作性表征。VLM在人类标注可操作性的2D图像上进行训练,无需机器人系统收集的训练数据。通过一种感知可操作性的数据合成流程,KALIE能基于少量人工收集的示例数据自动生成大规模高质量训练数据。实验表明,仅需50个示例数据点,KALIE即可稳健地解决包含未见物体的新操作任务。相比使用预训练VLM的基线方法,本方法始终表现出更优性能。

原文摘要 · Abstract (English)

Building generalist robotic systems involves effectively endowing robots with the capabilities to handle novel objects in an open-world setting. Inspired by the advances of large pre-trained models, we propose Keypoint Affordance Learning from Imagined Environments (KALIE), which adapts pre-trained Vision Language Models (VLMs) for robotic control in a scalable manner. Instead of directly producing motor commands, KALIE controls the robot by predicting point-based affordance representations based on natural language instructions and visual observations of the scene. The VLM is trained on 2D images with affordances labeled by humans, bypassing the need for training data collected on robotic systems. Through an affordance-aware data synthesis pipeline, KALIE automatically creates massive high-quality training data based on limited example data manually collected by humans. We demonstrate that KALIE can learn to robustly solve new manipulation tasks with unseen objects given only 50 example data points. Compared to baselines using pre-trained VLMs, our approach consistently achieves superior performance.

机器人控制视觉语言模型零样本操作数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。