给神经网络的‘世界模型’定义了可操作的标准,让研究者能客观判断是否真学到了世界表征。
What Does it Mean for a Neural Network to Learn a "World Model"?
- 通过线性探测思想,定义神经网络如何通过潜在状态空间表征世界生成过程。
- 提出检验标准,排除仅因数据或任务特性导致的虚假世界模型。
- 为实验验证提供统一语言,适合关注表征学习与模型可解释性的研究者。
我们提出一套精确标准,用于判断神经网络是否真正学习并使用了“世界模型”。目标是为常被模糊使用的术语赋予可操作的含义,促进实验研究的共同语言。本文聚焦于世界潜在状态空间的表示,暂不涉及动作建模。定义基于线性探测文献的思想,形式化了通过数据生成过程表征进行计算的机制。关键补充是若干检验条件,确保该“世界模型”并非神经网络数据或任务特性的平凡结果。
原文摘要 · Abstract (English)
We propose a set of precise criteria for saying a neural net learns and uses a "world model." The goal is to give an operational meaning to terms that are often used informally, in order to provide a common language for experimental investigation. We focus specifically on the idea of representing a latent "state space" of the world, leaving modeling the effect of actions to future work. Our definition is based on ideas from the linear probing literature, and formalizes the notion of a computation that factors through a representation of the data generation process. An essential addition to the definition is a set of conditions to check that such a "world model" is not a trivial consequence of the neural net's data or task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。