arXiv:2605.17204cs.ROcs.AI2026-05

将视觉语言动作模型的隐空间特征与机器人行为事件关联,提升可解释性。

Event-Grounded Sparse Autoencoders for Vision-Language-Action Policies

论文配图:Event-Grounded Sparse Autoencoders for Vision-Language-Action Policies
图 1 · 摘自论文原文
  • 用关键帧聚类锚定行为事件,连接特征与实际动作
  • 在仿真和真实机器人上验证了干预效果,提升策略可控性
  • 适合关注机器人决策透明度的研究者与工程师

视觉-语言-动作(VLA)策略将语言和视觉输入转化为机器人动作,其隐层表征直接影响闭环行为。然而,语言与视觉语言模型中的可解释性工具难以直接迁移到VLA:输出为机器人动作而非人类可读文本,且干预只能通过昂贵的闭环试运行验证。本文提出一种基于事件的可解释性流程,将稀疏自编码器(SAE)特征分析锚定在行为事件而非文本上下文。通过视觉、状态和时间线索对末端执行器关键帧进行聚类,将SAE特征与行为显著事件关联,并通过可选的视觉语言模型(VLM)标注链接至语义上下文。据我们所知,这是首个将SAE-based VLA分析扎根于闭环行为事件的工作。在两种仿真架构和一次真实机器人实验中,基于事件的排序表现出最强因果效应,且可迁移至π_{0.5}的连续动作片段。SAE是稀疏但不完美的干预基底:可用性随架构和干预位置变化,激进干预暴露出安全与可解释性局限。总体而言,基于事件的SAE分析为行为锚定的VLA可解释性提供了实用起点,推动未来研究向动作对齐坐标之外的特征、更细粒度闭环评估及高风险部署的安全干预方向发展。代码已开源。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) policies translate language and visual inputs into robot actions, where their hidden representations directly shape closed-loop behavior. However, mechanistic interpretability tools from language and vision-language models do not transfer cleanly to VLAs: outputs are robot actions rather than human-readable tokens, and interventions can only be tested via expensive closed-loop rollouts. We propose an event-grounded interpretability pipeline that anchors SAE feature analysis to behavioral events rather than text contexts. End-effector keyframes are clustered within each task using visual, state, and temporal cues, linking SAE features to behaviorally salient events and, via optional VLM annotations, to semantic context. To our knowledge, our pipeline is among the first to ground SAE-based VLA analysis in closed-loop behavioral events. Across two simulation architectures and a real-robot study, event-grounded ranking yields the strongest causal effects on OpenVLA and transfers to the continuous action chunks of $π_{0.5}$. SAE is a sparse but imperfect intervention basis: usability varies with architecture and intervention site, and aggressive intervention reveals safety and interpretability limits. Overall, event-grounded SAE analysis emerges as a practical starting point for behavior-anchored VLA interpretability, motivating future work on SAE features beyond action-aligned coordinates, finer-grained closed-loop evaluation, and safe interventions for high-stakes VLA deployments. Code is available at \url{https://github.com/xc-j/Event-SAE}.

VLA可解释性机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。