发现视觉语言动作模型能在线性读取任务进度,无需标签即可监控机器人执行状态。
Decoding Task Progress from VLA Representations

- 用线性探针从模型残差流中读取任务进度信号
- 该信号在未训练的PaliGemma中已存在,且跨任务泛化
- 适合用于部署后无监督异常检测,解释性强
视觉-语言-动作模型(VLAs)正快速走向通用操作策略的部署,但目前缺乏理解其内部表示或运行时监控的基本工具。借鉴机制可解释性思想,我们对π_{0.5}的残差流进行探测,发现任务进度(轨迹剩余时间的归一化值)可线性读取自激活值。该信号在预训练的PaliGemma骨干网络中即存在,无需任何机器人特定数据训练。单个线性探针在多提示数据上训练后,可泛化至未见任务,并在语言反事实条件下发生变化,但无法有效控制策略。这些特性使其直接适用于部署后模型的仪器化。我们利用该探针构建简单、无标签的分布外检测器,可识别任务停滞,性能媲美现有先进方法。结果表明,VLAs具有丰富且线性可读的内部语义表征,如任务进度,学习读取这些信号为监控部署的视觉运动策略提供了一条轻量、可解释的路径。
原文摘要 · Abstract (English)
Vision-language-action models (VLAs) are moving rapidly towards deployment as general-purpose manipulation policies, but we currently lack basic tools for understanding what these models represent internally or for monitoring them at runtime. Leveraging ideas from mechanistic interpretability, we probe the residual stream of $π_{0.5}$ and find that task progress, the normalized time remaining in a trajectory, is linearly readable from the activations. We find that this signal is present in the pretrained PaliGemma backbone prior to training on any robot-specific data. A single linear probe generalizes to unseen tasks and varies under language counterfactuals when trained on multi-prompt data, but does not enable meaningful steering of the policy. These properties make the signal directly useful for instrumenting deployed VLAs. We use the probe as a simple label-free OOD detector, which detects stalled task progress, and find it competitive with state-of-the-art methods. Our results suggest that VLAs have rich, linearly readable internal representations of semantic quantities like task progress, and that learning to read these signals offers a lightweight, interpretable path toward monitoring deployed visuomotor policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。