arXiv:2603.14523cs.CVcs.AI2026-03被引 2

让机器人像人一样边看边思考,提升复杂任务成功率

VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning

  • 把视觉信息当作可反复调用的推理动作,而非静态背景
  • 在LIBERO数据集上达97.5%成功率,长周期任务表现显著提升
  • 适合需要持续观察与动态决策的机器人操控场景

视觉-语言-动作(VLA)模型在具身智能中展现出潜力,但现有方法多依赖文本链式思维,将视觉输入视为静态上下文,限制了模型在长周期任务中主动回溯环境、解决歧义的能力。本文提出VLA-Thinker,一种基于图像思考的推理框架,将感知建模为可动态调用的推理动作。训练采用两阶段流程:第一阶段通过精选视觉链式思维数据进行监督微调,激活结构化推理与工具使用行为;第二阶段基于GRPO强化学习,对齐完整推理-动作轨迹与任务成功目标。在LIBERO和RoboTwin 2.0基准上的实验表明,VLA-Thinker显著提升操作性能,在LIBERO上达到97.5%的成功率,并在长周期机器人任务中取得显著进步。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown promising capabilities for embodied intelligence, but most existing approaches rely on text-based chain-of-thought reasoning where visual inputs are treated as static context. This limits the ability of the model to actively revisit the environment and resolve ambiguities during long-horizon tasks. We propose VLA-Thinker, a thinking-with-image reasoning framework that models perception as a dynamically invocable reasoning action. To train such a system, we introduce a two-stage training pipeline consisting of (1) an SFT cold-start phase with curated visual Chain-of-Thought data to activate structured reasoning and tool-use behaviors, and (2) GRPO-based reinforcement learning to align complete reasoning-action trajectories with task-level success. Extensive experiments on LIBERO and RoboTwin 2.0 benchmarks demonstrate that VLA-Thinker significantly improves manipulation performance, achieving 97.5% success rate on LIBERO and strong gains across long-horizon robotic tasks. Project and Codes: https://cywang735.github.io/VLA-Thinker/ .

机器人操控视觉推理强化学习具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。