arXiv:2603.16689cs.LG2026-03

用随机行走模拟世界,发现模型能自动提取世界结构的几何表示。

Predictive Statistics Shape Emergent World Representations of Grid Walkers

  • 用网格随机行走构建可控环境,预测依赖位置与剩余步数的统计量。
  • 变压器模型在首层注意力中提取世界状态,后续层映射到预测几何。
  • 该几何结构对不同约束通用,适合研究模型如何内化数据规律的人看。

下一步词预测器常表现出对潜在世界及其规则的内部表征。这些模型的概率特性暗示了世界结构与概率分布几何之间的深层联系。为更精确理解这一关联,我们采用一个最小化的随机过程作为受控设定:在二维格子上受约束的随机行走,必须在预设步数后抵达固定终点。最优预测仅依赖于由行者位置相对于目标及剩余时间窗决定的充分统计量;换言之,概率分布由网格几何参数化。我们在从该行走过程精确分布中采样的前缀上训练解码器仅变压器和循环网络,并通过测量各层隐藏激活与预测充分统计量之间的对齐度与线性可读性,进行比较。发现变压器计算分为两个阶段:首个注意力模块从输入中提取充分统计量,后续层将其转化为下一步的预测几何。在不同约束条件下,注意力后的表示具有普适性:一个共享的格子世界状态,可直接读取为世界模型,对应数据的预测几何。后续层则针对每种约束特化为下一步分布。循环网络虽达到贝叶斯最优损失,但未将此世界状态分离为独立阶段,表明世界模型几何也依赖于架构。尽管在玩具系统中演示,结果提示预测分布的几何是理解神经网络内化数据结构的有力视角。

原文摘要 · Abstract (English)

Next-token predictors often appear to develop internal representations of the latent world and its rules. The probabilistic nature of these models suggests a deep connection between the structure of the world and the geometry of probability distributions. In order to understand this link more precisely, we use a minimal stochastic process as a controlled setting: constrained random walks on a two-dimensional lattice that must reach a fixed endpoint after a predetermined number of steps. Optimal prediction of this process solely depends on a sufficient vector determined by the walker's position relative to the target and the remaining time horizon; in other words, the probability distributions are parametrized by the world's grid geometry. We train decoder-only transformers and recurrent networks on prefixes sampled from the exact distribution of these walks and compare their hidden activations to sufficient statistics of prediction, by measuring alignment and linear readability across layers. We find that the transformer's computation factors into two stages: the first attention block extracts the sufficient statistic from the input, and later layers transform it into the next-step predictive geometry. Across constraint variants the post-attention representation is universal: a shared world-state of the lattice that can be read directly as a world model, traced to the predictive geometry of the data. Later layers then specialize it to each variant's next-step distribution. Recurrent networks reach the same Bayes-optimal loss but do not isolate this world-state as a separate stage, showing that the world-model geometry also depends on architecture. Although demonstrated in a toy system, the results suggest that the geometry of the predictive distribution is a useful lens on how neural networks internalize the structure of their data.

世界模型概率几何注意力机制生成建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。