arXiv:2608.05738cs.RO2026-08

让视觉语言动作模型通过上下文学习理解语言,提升真实机器人操作性能。

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use

论文配图:In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use
图 1 · 摘自论文原文
  • 用上下文后训练注入感知信息,仅监督动作,避免语言与动作目标冲突。
  • 在多个仿真和真实机器人任务中表现优于基于思维链的方法。
  • 通过智能工具主动获取信息,实现对未见过语言的精准理解。

视觉-语言-动作(VLA)模型已成为通用操作的主流方法,但普遍采用行为克隆训练:策略根据静态图像和固定指令模仿专家动作片段。一种自然改进是引入文本思维链(CoT)进行显式推理。我们通过实验和分析表明,自由格式的文本思维链会损害底层控制:产生的推理缺乏实体依据,延迟破坏闭环时序,且推理与动作令牌优化目标相冲突,导致策略更擅长叙述而非执行。我们认为VLA真正需要的不是生成语言的能力,而是理解具身语言的能力。为此,我们提出 extbf{ hiswork{}}框架,通过(i)上下文后训练,将感知证据作为结构化上下文注入,仅对动作进行监督;(ii)代理式工具接口,使策略调用开放词汇检测器、单目深度和视觉-语言模型,主动获取任务相关信息。数据引擎生成多样、改写、基于证据的空间描述,使策略学会理解从未见过的表述。在RoboCasa-GR1、SimplerEnv、LIBERO仿真基准及8个真实机器人操作任务中,该方法在性能和效率上均优于对比的思维链方法,配置一致下达到最先进水平。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have become the dominant recipe for generalist manipulation, yet they are almost universally trained by behavior cloning: a policy imitates expert action chunks conditioned on a static image and a fixed instruction. A natural remedy is to inject explicit reasoning through textual chain-of-thought (CoT). We show, both empirically and analytically, that free-form textual CoT degrades low-level control: the reasoning it produces is ungrounded, its latency breaks closed-loop timing, and, crucially, the reasoning and action tokens are optimized against conflicting objectives so that the policy learns to narrate rather than to act. We argue that what a VLA needs is not the ability to generate language, but the ability to consume grounded language. To this end we introduce \textbf{\ourmethod{}}, a framework that endows a VLA with language competence through (i) in-context post-training, in which perceptual evidence is injected as structured context and the model is supervised only on actions, and (ii) an agentic tool-use interface, in which the policy queries open-vocabulary detectors, monocular depth, and a vision--language model to actively acquire task-relevant information. Rather than emitting a single templated caption, our data engine produces diverse, paraphrased, and evidence-conditioned spatial descriptions, so that the policy learns to interpret language it has never seen verbatim. Across the RoboCasa-GR1, SimplerEnv, and LIBERO simulation benchmarks, together with 8 real-world robot manipulation tasks, our method consistently achieves SOTA results in both performance and efficiency when compared with CoT-based approaches under matched configurations.

视觉语言动作上下文学习机器人操作智能体工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。