发现顶尖视觉语言动作模型暗藏世界模型,训练越久越明显。
Emergent World Representations in OpenVLA
- 用嵌入运算探测模型内部状态转移规律
- 实验证明模型能准确预测状态变化,优于基础嵌入方法
- 适合研究具身智能与模型内隐表征的学者参考
基于策略的强化学习训练的视觉语言动作模型(VLAs)能在不显式建模环境动态的情况下实现复杂行为。然而,这些模型是否隐式学习了世界模型仍是未知。本文提出一种实验方法,通过状态表示的嵌入算术,探测当前最先进的VLA模型OpenVLA是否包含状态转移的潜在知识。具体地,测量序列环境状态嵌入间的差异,并检验该转移向量能否从中间层激活中恢复。利用在线性和非线性探针在多层激活上训练,结果表明其对状态转移具有统计显著的预测能力,且优于基线嵌入方法,说明OpenVLA确实编码了内部世界模型(而非探针自身学习转移)。我们还分析了OpenVLA早期检查点,发现世界模型随训练进程逐渐显现。最后,提出了一个基于稀疏自编码器(SAEs)的分析管道,用于进一步解码模型的世界模型。
原文摘要 · Abstract (English)
Vision Language Action models (VLAs) trained with policy-based reinforcement learning (RL) encode complex behaviors without explicitly modeling environmental dynamics. However, it remains unclear whether VLAs implicitly learn world models, a hallmark of model-based RL. We propose an experimental methodology using embedding arithmetic on state representations to probe whether OpenVLA, the current state of the art in VLAs, contains latent knowledge of state transitions. Specifically, we measure the difference between embeddings of sequential environment states and test whether this transition vector is recoverable from intermediate model activations. Using linear and non linear probes trained on activations across layers, we find statistically significant predictive ability on state transitions exceeding baselines (embeddings), indicating that OpenVLA encodes an internal world model (as opposed to the probes learning the state transitions). We investigate the predictive ability of an earlier checkpoint of OpenVLA, and uncover hints that the world model emerges as training progresses. Finally, we outline a pipeline leveraging Sparse Autoencoders (SAEs) to analyze OpenVLA's world model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。