用物体和机器人自身信息做视觉令牌,大幅降低计算量却保持性能
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
- 以物体和机器人自身为中心生成视觉令牌,引入对象感知先验
- 仅需少量视觉令牌即可达到与原模型相当的性能
- 训练速度比OpenVLA快一倍以上,适用于真实抓取任务
视觉-语言-动作(VLA)模型通过复用大规模预训练视觉-语言模型(VLM)来输出机器人动作,为规模化学习机器人操作提供了重要路径。然而,将VLM适配到机器人领域带来过高的计算开销,我们归因于视觉输入的令牌化方式。本文提出Oat-VLA,一种面向VLA的对象-机器人中心令牌化方法。基于对象中心表示学习的洞察,该方法引入场景物体及机器人自身视觉信息的归纳偏置。结果表明,Oat-VLA可将视觉令牌数量大幅减少至仅几枚,同时保持性能。在LIBERO基准上,Oat-VLA的收敛速度至少是OpenVLA的两倍,并在多样化的现实世界抓取放置任务中表现更优。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models offer a pivotal approach to learning robotic manipulation at scale by repurposing large pre-trained Vision-Language-Models (VLM) to output robotic actions. However, adapting VLMs for robotic domains comes with an unnecessarily high computational cost, which we attribute to the tokenization scheme of visual inputs. In this work, we aim to enable efficient VLA training by proposing Oat-VLA, an Object-Agent-centric Tokenization for VLAs. Building on the insights of object-centric representation learning, our method introduces an inductive bias towards scene objects and the agent's own visual information. As a result, we find that Oat-VLA can drastically reduce the number of visual tokens to just a few tokens without sacrificing performance. We reveal that Oat-VLA converges at least twice as fast as OpenVLA on the LIBERO suite, as well as outperform OpenVLA in diverse real-world pick and place tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。