研究冻结的视觉语言动作模型如何编码、使用和操控历史信息。
Present but Not Remembered: Auditing How Frozen VLAs Encode, Deploy, and Steer Visual History
- 通过分层线性探针与因果干预分析历史信息在模型中的表征。
- 历史信息虽被编码但多为当前帧冗余副本,仅在图像严重遮挡时才被调用。
- 模型部署历史的方式决定可操控性,而非是否存储历史。
一个冻结的视觉-语言-动作模型(VLA)在每个决策步骤接收最新观测,但以往工作侧重于增加记忆,而非探究现有历史如何被表示与利用。本文通过分层线性探针与因果交换干预,在来自两个架构家族的三个VLA上研究时间维度。发现三重分离:第一,过去帧内容在整个网络中仍可线性解码;第二,超越当前帧的独特历史信息几乎不存在,表明存储的历史主要是当前帧的冗余拷贝;第三,历史仅在当前帧严重退化时被因果调用,而动作读出则随网络层级逐步降低对历史的依赖。尽管所有模型编码历史方式相似,其部署策略却不同:相同遮挡下,一种架构逐渐依赖历史作为后备,另一种则较少依赖。我们提出一种无需训练的时间部署审计方法,可区分这两种模式。在后备模式下,重新注入历史既无法修复遮挡也无法消除动作歧义,证实其表征的冗余性;而在另一模式下,相同干预能可靠引导预测动作向源历史靠拢。结果表明,可控性取决于历史的部署方式,而非是否编码。VLAs并非遗忘过去,而是未能将其表示为区别于当前的信息。未来记忆增强应注入过去独有的信息,而非简单增加历史数据。
原文摘要 · Abstract (English)
A frozen vision-language-action model (VLA) receives recent observations at every decision step, yet prior work has focused on adding memory rather than asking how existing history is represented and used. We study this temporal axis using layer-resolved linear probing and causal interchange interventions across three VLAs from two architecture families. We find a three-part dissociation. First, past-frame content remains linearly decodable throughout the network. Second, information unique to history beyond the current frame is nearly absent, indicating that stored history is largely a redundant copy of the present. Third, history is causally deployed only when the current frame is heavily degraded, while the action readout progressively loses dependence on history through the network. Although all models encode history similarly, their deployment strategies differ: under the same occlusion, one architecture increasingly relies on history as a fallback, whereas the other relies on it less. We further introduce a training-free temporal deployment audit that distinguishes these regimes. In the fallback regime, re-injecting history neither repairs occlusion nor disambiguates actions, confirming the redundancy of the stored representation. In the other regime, the same intervention reliably steers the predicted action toward the donor history. These results show that steerability depends on how history is deployed rather than whether it is encoded. VLAs do not forget the past; they largely fail to represent it as information distinct from the present. Our findings suggest that future memory augmentation should inject information unique to the past rather than simply more history.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。