arXiv:2603.19183cs.RO2026-03被引 10

用稀疏自编码器揭示视觉语言动作模型的可解释特征

Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models

  • 用稀疏自编码器分析模型隐藏层激活,发现可解释特征
  • 识别出跨任务通用的动作基元和语义概念,可被因果调控
  • 适用于研究机器人泛化机制或做可控行为干预

视觉-语言-动作(VLA)模型在通用机器人操作中展现出巨大潜力,但其跨物体、场景和指令的泛化机制仍不清晰。为探究内部表征,我们在VLA模型的隐藏层激活上训练稀疏自编码器(SAE),学习激活空间中的稀疏字典,揭示出与可解释方向对应的特征。我们识别出对应于运动基元和语义概念的SAE特征,包括跨任务通用且可因果调控的特征。提出一种指标以区分通用迁移基元与特定任务记忆。在LIBERO仿真基准和真实DROID硬件上验证,放大通用语义特征可诱发符合语义的行为,而移除则导致性能下降。进一步展示可通过调控实现非提示方向的行为控制。结果表明VLA能学习到连接感知、语言与动作的可重用内部特征。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have emerged as a promising approach for general-purpose robot manipulation. However, little research has mechanistically explored when and why they generalize across objects, scenes, and instructions. To probe internal representations, we train Sparse Autoencoders (SAEs) on the VLA's hidden-layer activations. SAEs learn sparse dictionaries over model activations, often revealing features that correspond to interpretable directions in the model's representation space. We identify SAE features corresponding to motion primitives and semantic concepts, including features that are general across episodes and causally steerable. We propose a metric to categorize features as general transferable primitives or episode-specific memorizations, offering a promising glimpse towards VLA generalization. We validate these findings through steering experiments on both the LIBERO simulation benchmark and on real-world DROID hardware. We find that amplifying general and semantic features induces behaviors consistent with their meanings, whereas ablating them destroys model performance. Furthermore, we demonstrate steering as a way to control behavior in unpromptable directions. Together, these results provide mechanistic evidence that VLAs can learn reusable internal features linking perception, language, and action across tasks and scenes. Our project page is located at https://drvla.github.io

VLA模型可解释性动作基元特征操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。