arXiv:2605.17144cs.ROcs.AI2026-05被引 2

通过隐状态引导,让视觉语言动作模型在机器人任务中更稳定成功。

Contrastive Conceptor Activation Steering (COAST): Unlocking Vision-Language-Action Models through Hidden States

论文配图:Contrastive Conceptor Activation Steering (COAST): Unlocking Vision-Language-Action Models through Hidden States
图 1 · 摘自论文原文
  • 用概念器识别成功与失败的隐空间特征子集
  • 仿真和真实机器人任务成功率分别提升20%和40%以上
  • 无需重新训练,可跨任务复用失败模式结构

视觉语言动作(VLA)模型利用大规模视觉语言模型的感知先验,但在实际机器人任务中仍表现脆弱。本文提出对比概念器激活引导(COAST),基于“概念器”这一线性算子,从少量成功与失败轨迹中识别目标任务的关键隐空间子集。推理时,该方法将VLA的隐状态引导至成功子空间,显著提升任务成功率。在三种不同架构的神经策略(流匹配VLA、自回归VLA、Diffusion Policy)上,仿真任务成功率提升超20%,真实机器人任务提升超40%。分析显示,失败模式在不同任务间具有高度结构相似性,而成功表示则高度任务特异。当任务共享相似失败模式时,已拟合的概念器可直接用于新任务,无需重训练。结果表明,现有VLA在隐状态中仍保留丰富任务相关知识,通过引导其残差流至任务相关子空间,可缓解动作解码瓶颈。COAST提供了一种无需训练、轻量高效的路径,释放模型内部的潜在能力。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models leverage powerful perceptual priors from web-scale Vision-Language Model (VLM) pre-training, yet they remain surprisingly brittle in practice, frequently failing at simple robotic tasks. To mitigate this, we propose Contrastive Conceptor Activation Steering (COAST). COAST builds on the notion of a "conceptor", a linear operator that soft-projects data into the principal components of a target distribution. COAST uses conceptors to identify success-critical subspaces for a target robotic task from a few examples of success and failure rollouts. At inference time, it steers VLA latents into these identified success subspaces to improve task outcomes. Across three architecturally distinct neural policies (flow-matching VLA, autoregressive VLA, and Diffusion Policy), COAST improves absolute mean simulation and real-robot task success rate by over 20 and 40% respectively. The activation subspace geometry reveals that failure modes share substantial structure across tasks while success representations remain largely task-specific. When tasks share similar failure modes, this structure enables previously fitted conceptors to improve performance on new tasks without refitting. Ultimately, our results suggest that current VLAs retain substantial task-relevant knowledge in their latent representations, and that the action expert's decoding bottleneck could be mitigated by steering its residual stream toward task-relevant subspaces. COAST provides a lightweight, training-free path to unlocking these latent capabilities by steering the model towards its own "success" distributions.

视觉语言动作隐空间引导机器人控制零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。