用视觉指令提升机器人抓取任务的泛化能力
VIP: Vision Instructed Pre-training for Robotic Manipulation
- 用未来图像+稀疏点流作为视觉指令指导机器人动作
- 在真实与仿真环境中显著提升多任务完成率
- 适合需要高精度目标识别的机器人操控场景
机器人操控中扩大训练数据的效果仍受限。主要挑战在于任务多样,若目标不明确,策略容易混淆。现有方法依赖文本指令描述目标,但我们发现当前机器人数据无法有效训练出理解文本指令的策略,而视觉信息更易理解。因此,我们提出使用视觉指令明确目标。直接做法是训练策略预测连接当前观测与未来图像的中间动作。然而单张未来图像信息不足。为此,我们引入稀疏点流提供更详细信息。基于真实与仿真环境设计了多种任务评估所提视觉指令预训练(VIP)方法。结果表明,VIP显著提升了多样化任务的表现,衍生策略可完成如‘打开紧封瓶盖’等复杂任务。
原文摘要 · Abstract (English)
The effectiveness of scaling up training data in robotic manipulation is still limited. A primary challenge in manipulation is the tasks are diverse, and the trained policy would be confused if the task targets are not specified clearly. Existing works primarily rely on text instruction to describe targets. However, we reveal that current robotic data cannot train policies to understand text instruction effectively, and vision is much more comprehensible. Therefore, we introduce utilizing vision instruction to specify targets. A straightforward implementation is training a policy to predict the intermediate actions linking the current observation and a future image. Nevertheless, a single future image does not describe the task target in insufficient detail. To handle this problem, we propose to use sparse point flows to provide more detailed information. Extensive tasks are designed based on real and simulated environments to evaluate the effectiveness of our vision instructed pre-training (VIP) method. The results indicate VIP improves the performance on diverse tasks significantly, and the derived policy can complete competitive tasks like ``opening the lid of a tightly sealed bottle''.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。