arXiv:2603.19233cs.RO2026-03中稿 · ICLR被引 9

揭秘视觉语言动作模型如何把看到的和说的变成动作

Not All Features Are Created Equal: A Mechanistic Study of Vision-Language-Action Models

论文配图:Not All Features Are Created Equal: A Mechanistic Study of Vision-Language-Action Models
图 1 · 摘自论文原文
  • 用激活注入等方法分析六种模型,发现视觉信息主导动作生成
  • 语言只在任务模糊时起作用,94%到10%的行为差异体现其关键性
  • 首次揭示动作程序与目标语义分属不同激活空间,适合机器人研究者

视觉-语言-动作(VLA)模型将感知、语言和运动控制整合于单一架构,但其如何将多模态输入转化为动作仍不清晰。我们对六种模型(参数量80M至7B)在四个基准上超过39.4万次推演展开研究,采用激活注入、稀疏自编码器(SAEs)和线性探测。结果显示,视觉路径主导动作生成:在空提示场景中注入基线激活即可恢复近似行为;跨任务注入可引导机器人至源任务位置(X-VLA 99.8%轨迹对齐)。语言敏感性取决于任务结构:当视觉唯一确定任务时语言被忽略;当多个目标共享场景时语言至关重要(X-VLA libero_goal任务错误提示下成功率从94%降至10%,而libero_object任务维持60–100%)。三种多路径架构(pizhalf, SmolVLA, GR00T)中,专家路径编码运动程序,视觉语言模型路径编码目标语义(专家注入导致行为偏移2倍),子空间注入证实二者分离。多数模型需逐标记处理以保动作精度,但平均池化提升X-VLA表现。对比识别提取出82+操作概念,因果消融显示零效应率在28%–92%之间,不受表征宽度影响。我们发布行动图谱(Action Atlas,https://action-atlas.com)供交互探索所有六种模型的表示。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models combine perception, language, and motor control in a single architecture, yet how they translate multimodal inputs into actions remains poorly understood. We apply activation injection, sparse autoencoders (SAEs), and linear probes to six models spanning 80M--7B parameters across 394,000+ rollout episodes on four benchmarks. The visual pathway dominates action generation across all architectures: injecting baseline activations into null-prompt episodes recovers near-identical behavior, while cross-task injection steers robots toward source-task positions (99.8\% of X-VLA episodes align with the source trajectory), exposing spatially bound motor programs tied to scene coordinates rather than abstract task representations. Language sensitivity depends on task structure, not model design: when visual context uniquely specifies the task, language is ignored; when multiple goals share a scene, language becomes essential (X-VLA \texttt{libero\_goal}: 94\%$\to$10\% under wrong prompts vs.\ \texttt{libero\_object}: 60--100\% regardless). In all three multi-pathway architectures (\pizhalf{}, SmolVLA, GR00T), expert pathways encode motor programs while VLM pathways encode goal semantics ($2\times$ greater behavioral displacement from expert injection), and subspace injection confirms these occupy separable activation subspaces. Per-token SAE processing is essential for action fidelity on most architectures, though mean-pooling improves fidelity on X-VLA. Contrastive identification recovers 82+ manipulation concepts, and causal ablation reveals sensitivity spanning 28--92\% zero-effect rates independent of representation width. We release \textbf{Action Atlas} (https://action-atlas.com) for interactive exploration of VLA representations across all six models.

视觉语言动作神经机制机器人控制模型解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。