arXiv:2509.00328cs.ROcs.LG2025-09被引 26

通过内部表示解析视觉语言动作模型,实现无需训练的实时行为调控。

Mechanistic interpretability for steering vision-language-action models

  • 将变压器层激活投影到词嵌入基,识别影响动作选择的稀疏语义方向。
  • 在模拟和真实机器人上实现零样本行为控制,无需微调或奖励信号。
  • 为机器人基础模型提供可解释、可干预的新型控制范式,适合研究可信赖智能体。

视觉-语言-动作(VLA)模型是实现通用具身智能体的有前景路径,能快速适应新任务、模态和环境。然而,现有解释与调控方法远不及传统机器人系统,后者基于显式的运动学、动力学和控制模型。这一机制性理解的缺失是部署学习策略于真实机器人中的核心挑战,因鲁棒性和可解释性至关重要。受大语言模型机制可解释性进展启发,我们首次提出通过内部表征解释和调控VLA的框架,实现在推理时直接干预模型行为。我们将变压器层中的前馈激活投影至词嵌入基,识别出与动作选择因果关联的稀疏语义方向(如速度、方向)。基于此,提出一种通用激活调控方法,可在无微调、无奖励信号、无环境交互下实时调节行为。我们在两个开源VLA模型Pi0和OpenVLA上评估,展示了在模拟(LIBERO)和物理机器人(UR5)上的零样本行为控制。本工作证明,具身VLA中可解释组件可系统化用于控制,建立了一种透明且可调控的基础模型新范式。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models are a promising path to realizing generalist embodied agents that can quickly adapt to new tasks, modalities, and environments. However, methods for interpreting and steering VLAs fall far short of classical robotics pipelines, which are grounded in explicit models of kinematics, dynamics, and control. This lack of mechanistic insight is a central challenge for deploying learned policies in real-world robotics, where robustness and explainability are critical. Motivated by advances in mechanistic interpretability for large language models, we introduce the first framework for interpreting and steering VLAs via their internal representations, enabling direct intervention in model behavior at inference time. We project feedforward activations within transformer layers onto the token embedding basis, identifying sparse semantic directions - such as speed and direction - that are causally linked to action selection. Leveraging these findings, we introduce a general-purpose activation steering method that modulates behavior in real time, without fine-tuning, reward signals, or environment interaction. We evaluate this method on two recent open-source VLAs, Pi0 and OpenVLA, and demonstrate zero-shot behavioral control in simulation (LIBERO) and on a physical robot (UR5). This work demonstrates that interpretable components of embodied VLAs can be systematically harnessed for control - establishing a new paradigm for transparent and steerable foundation models in robotics.

机器人控制可解释性模型干预具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。