让机器人具备像人一样看懂数学题并动手解决的能力。
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
- 用专家混合架构和两阶段训练,保留视觉语言模型的通用知识
- 无需专门训练就能解数学题、识字卡,准确率超现有方法
- 适合想构建强推理能力机器人的研究者与开发者
视觉-语言-动作(VLA)模型已成为机器人领域的下一代模型。尽管利用了强大的预训练视觉语言模型(VLM),现有端到端VLA系统在微调过程中常丧失关键能力。我们认为,通用VLA模型应保留并扩展VLM的核心能力:1)开放世界具身推理——继承VLM的知识,能识别任何VLM可识别的内容,解决数学问题,具备视觉空间智能;2)推理跟随——将开放世界的推理有效转化为机器人可执行的动作。本文提出ChatVLA-2,一种新型专家混合VLA模型,搭配专用的两阶段训练流程,旨在保留VLM原始优势的同时实现可行动作推理。为验证方法,我们设计了一个数学匹配任务:机器人需解读白板上的数学题,并从桌面上选取对应数字卡片解答方程。令人惊讶的是,该方法虽未显式训练数学或OCR能力,仍表现出卓越的数学推理与光学字符识别性能。此外,实验表明其具备强空间推理能力,可理解涉及从未见过物体的全新方向指令。整体上,该方法在推理与理解能力上显著优于当前最先进的模仿学习方法,如OpenVLA、DexVLA和pi-zero。本工作标志着向具备强大推理能力的通用机器人基础模型迈出了重要一步。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), existing end-to-end VLA systems often lose key capabilities during fine-tuning as the model adapts to specific robotic tasks. We argue that a generalizable VLA model should retain and expand upon the VLM's core competencies: 1) Open-world embodied reasoning - the VLA should inherit the knowledge from VLM, i.e., recognize anything that the VLM can recognize, be capable of solving math problems, and possess visual-spatial intelligence, 2) Reasoning following - effectively translating the open-world reasoning into actionable steps for the robot. In this work, we introduce ChatVLA-2, a novel mixture-of-expert VLA model coupled with a specialized two-stage training pipeline designed to preserve the VLM's original strengths while enabling actionable reasoning. To validate our approach, we design a math-matching task wherein a robot interprets math problems written on a whiteboard and picks corresponding number cards from a table to solve equations. Remarkably, our method exhibits exceptional mathematical reasoning and OCR capabilities, despite these abilities not being explicitly trained within the VLA. Furthermore, we demonstrate that the VLA possesses strong spatial reasoning skills, enabling it to interpret novel directional instructions involving previously unseen objects. Overall, our method showcases reasoning and comprehension abilities that significantly surpass state-of-the-art imitation learning methods such as OpenVLA, DexVLA, and pi-zero. This work represents a substantial advancement toward developing truly generalizable robotic foundation models endowed with robust reasoning capacities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。