用视觉思维链统一感知与动作,让机器人更懂环境、更准执行。
Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation
- 通过隐式视觉思维链,将图像预测与机械臂动作联合生成。
- 在多个仿真和真实任务中,成功率最高提升14.5%,平均达80.5%。
- 适合做通用机器人操作模型,尤其擅长复杂空间任务。
基于思维链(CoT)的视觉-语言-动作(VLA)模型在通用机器人代理中取得显著进展,得益于其强大的感知理解能力。然而,纯文本思维链难以捕捉复杂空间环境中的细节,因此引入视觉先验成为关键策略。现有方法面临两大挑战:一是视觉观测与低级动作之间的模态鸿沟;二是视觉预测与动作生成目标冲突导致训练不稳定。为此,我们提出视觉融合轨迹对齐(VITA)框架,学习视觉与动作的共享离散潜在空间,实现感知与运动控制的联合建模。VITA引入隐式视觉思维链:自回归生成的令牌同时解码为未来帧预测和机器人动作,将视觉动态作为运动规划的归纳偏置。在仿真与真实世界环境中的大量实验表明,VITA在CALVIN、LIBERO和SimplerEnv上分别超越基线14.5%、9.6%和12.1%。此外,其在六项真实任务中平均成功率达80.5%,展现出作为通用机器人操作模型的巨大潜力。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models built upon Chain-of-Thought (CoT) have achieved remarkable success in advancing general-purpose robotic agents, owing to its significant perceptual comprehension. Recently, since text-only CoT struggles to adequately capture scene details in complex spatial environments, a highly promising strategy involves leveraging visual priors to guide robotic action generation. Nevertheless, these strategies face two inherent challenges: (i) a modality gap between visual observations and low-level actions, and (ii) unstable training due to competing objectives between visual prediction and action generation. To address these challenges, we propose a Vision-Integrated Trajectory Alignment (VITA) framework that learns a shared discrete latent space for vision and action, enabling joint modeling of perception and motor control. VITA introduces a implicit visual CoT: autoregressively generated tokens is simultaneously decoded into future frames predictions and robot actions, thereby internalizing visual dynamics as an inductive bias for motion planning. Extensive experiments on simulated and real-world environments demonstrate state-of-the-art performance. VITA improves 14.5\%, 9.6\% and 12.1\% over existing baselines on CALVIN, LIBERO and SimplerEnv. Furthermore, VITA attains an average success rate of 80.5\% across six real-world tasks, demonstrating its potential as a generalist robotic manipulation model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。