通过自监督学习视觉与文本关联,实现无需动作标注的机器人抓取表征。
LaVA-Man: Learning Visual Action Representations for Robot Manipulation
- 用图文条件重建遮蔽目标图像,自监督学习视觉-动作表征。
- 仅需少量示范即可在5个基准上超越现有方法,真实机器人验证有效。
- 新数据集支持180类物体泛化,适合研究语言引导机器人操作的学者。
视觉-文本理解对语言引导的机器人操作至关重要。现有方法通常分两步:先用预训练视觉-语言模型计算视觉观测与文本指令的相似性,再训练模型将该相似性映射为机器人动作。这种两阶段方式限制了模型捕捉视觉与文本间深层关系的能力,导致操作精度下降。本文提出一种自监督预训练任务:在输入图像和文本指令条件下,重建被遮蔽的目标图像。该方法使模型无需机器人动作标注即可学习视觉-动作表征,并可在仅有少量示范的情况下微调用于实际操作任务。我们还引入了新的「Omni-Object Pick-and-Place」数据集,包含180类物体、3200个实例及对应文本指令,支持多样化物体先验学习,并全面评估模型在不同物体实例间的泛化能力。在五个基准(含模拟与真实机器人)上的实验结果表明,本方法显著优于先前方法。
原文摘要 · Abstract (English)
Visual-textual understanding is essential for language-guided robot manipulation. Recent works leverage pre-trained vision-language models to measure the similarity between encoded visual observations and textual instructions, and then train a model to map this similarity to robot actions. However, this two-step approach limits the model to capture the relationship between visual observations and textual instructions, leading to reduced precision in manipulation tasks. We propose to learn visual-textual associations through a self-supervised pretext task: reconstructing a masked goal image conditioned on an input image and textual instructions. This formulation allows the model to learn visual-action representations without robot action supervision. The learned representations can then be fine-tuned for manipulation tasks with only a few demonstrations. We also introduce the \textit{Omni-Object Pick-and-Place} dataset, which consists of annotated robot tabletop manipulation episodes, including 180 object classes and 3,200 instances with corresponding textual instructions. This dataset enables the model to acquire diverse object priors and allows for a more comprehensive evaluation of its generalisation capability across object instances. Experimental results on the five benchmarks, including both simulated and real-robot validations, demonstrate that our method outperforms prior art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。