arXiv:2411.19650cs.ROcs.AI2024-11被引 431

让机器人更懂指令,能灵活应对新环境和新物体。

CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation

  • 将视觉语言模型输出与专用动作模块结合,用扩散模型生成动作序列。
  • 在仿真和真实机器人上成功率超同类模型35%以上,远超55B大模型。
  • 适配新机器人、新物品和新场景能力强,适合通用机械臂任务部署。

大型视觉-语言-动作(VLA)模型的进展显著提升了机器人在语言引导下的任务执行能力与未见场景泛化能力。尽管现有基于预训练视觉-语言模型(VLM)的VLA展现出一定泛化性,其任务成功率仍不理想。本文提出一种新型高级VLA架构,源自VLM但采用组件化设计,动作模块由VLM输出条件驱动。我们系统研究了动作模块设计,发现使用扩散动作变换器建模动作序列可显著提升性能并具备良好扩展性。在5种仿真与真实机器人上的综合实验与消融分析表明,本模型不仅大幅超越现有VLA,且对新机器人、未见物体和背景具有出色适应性。在仿真中成功率超过同规模模型OpenVLA(7B)35%以上,在真实机器人实验中高出55%;在仿真中亦比更大模型RT-2-X(55B)高18个百分点。代码与模型详见项目页:https://cogact.github.io/

原文摘要 · Abstract (English)

The advancement of large Vision-Language-Action (VLA) models has significantly improved robotic manipulation in terms of language-guided task execution and generalization to unseen scenarios. While existing VLAs adapted from pretrained large Vision-Language-Models (VLM) have demonstrated promising generalizability, their task performance is still unsatisfactory as indicated by the low tasks success rates in different environments. In this paper, we present a new advanced VLA architecture derived from VLM. Unlike previous works that directly repurpose VLM for action prediction by simple action quantization, we propose a omponentized VLA architecture that has a specialized action module conditioned on VLM output. We systematically study the design of the action module and demonstrates the strong performance enhancement with diffusion action transformers for action sequence modeling, as well as their favorable scaling behaviors. We also conduct comprehensive experiments and ablation studies to evaluate the efficacy of our models with varied designs. The evaluation on 5 robot embodiments in simulation and real work shows that our model not only significantly surpasses existing VLAs in task performance and but also exhibits remarkable adaptation to new robots and generalization to unseen objects and backgrounds. It exceeds the average success rates of OpenVLA which has similar model size (7B) with ours by over 35% in simulated evaluation and 55% in real robot experiments. It also outperforms the large RT-2-X model (55B) by 18% absolute success rates in simulation. Code and models can be found on our project page (https://cogact.github.io/).

机器人操控视觉语言动作扩散模型通用智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。