让视觉语言动作模型像人一样分步思考,提升复杂操作能力
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

- 用未来图像作为视觉思维链,分步规划动作
- 实测比现有最佳模型高17%的现实任务表现
- 适合需要长期规划的机器人操作研究者
视觉语言动作模型(VLAs)在利用预训练视觉语言模型和多样机器人示范学习通用感知运动控制方面展现出潜力。尽管该范式有效整合了机器人与非机器人数据,但当前VLAs主要依赖直接输入输出映射,缺乏复杂操作任务所需的中间推理步骤,因而缺乏时间规划或推理能力。本文提出CoT-VLA,通过自回归预测未来图像帧作为视觉目标,再生成短序列动作以达成这些目标,引入显式的视觉思维链(CoT)推理。CoT-VLA是当前最先进的70亿参数VLA,能理解并生成视觉与动作标记。实验结果表明,CoT-VLA在真实世界操作任务中比最先进模型高出17%,在仿真基准测试中提升6%。
原文摘要 · Abstract (English)
Vision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor control. While this paradigm effectively utilizes large-scale data from both robotic and non-robotic sources, current VLAs primarily focus on direct input--output mappings, lacking the intermediate reasoning steps crucial for complex manipulation tasks. As a result, existing VLAs lack temporal planning or reasoning capabilities. In this paper, we introduce a method that incorporates explicit visual chain-of-thought (CoT) reasoning into vision-language-action models (VLAs) by predicting future image frames autoregressively as visual goals before generating a short action sequence to achieve these goals. We introduce CoT-VLA, a state-of-the-art 7B VLA that can understand and generate visual and action tokens. Our experimental results demonstrate that CoT-VLA achieves strong performance, outperforming the state-of-the-art VLA model by 17% in real-world manipulation tasks and 6% in simulation benchmarks. Project website: https://cot-vla.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。