arXiv:2603.04317cs.CLcs.AI2026-03被引 1

静态词嵌入自带地理时间结构,无需世界模型。

World Properties without World Models: Recovering Spatial and Temporal Structure from Co-occurrence Statistics in Static Word Embeddings

  • 用词共现嵌入直接恢复地理与时间信息,不依赖复杂模型。
  • 城市坐标可解释方差达71%-87%,出生年份达48%-52%。
  • 关键在国家名和气候词汇等语义梯度,适合对语言结构感兴趣者。

近期研究将大语言模型隐藏状态中可线性恢复的地理与时间变量视为世界表征的证据。我们检验了一个更简单的可能性:相关结构可能早已潜藏于文本本身。对静态共现嵌入(GloVe与Word2Vec)应用相同的岭回归探针,发现显著可恢复的地理信号和较弱但可靠的时序信号,城市坐标在留出数据上的R²为0.71-0.87,历史出生年份为0.48-0.52。语义邻域分析与子空间消融显示,这些信号强烈依赖可解释的词汇梯度,尤其是国家名称与气候相关词汇。结果表明,普通词共现嵌入比通常认为的保留了更丰富的空间、时间与环境结构,揭示了静态嵌入仅从文本中便具备强大世界结构表征能力。因此,线性可恢复性本身不足以证明超越文本的表征跃迁。

原文摘要 · Abstract (English)

Recent work interprets the linear recoverability of geographic and temporal variables from large language model (LLM) hidden states as evidence for world-like internal representations. We test a simpler possibility: that much of the relevant structure is already latent in text itself. Applying the same class of ridge regression probes to static co-occurrence-based embeddings (GloVe and Word2Vec), we find substantial recoverable geographic signal and weaker but reliable temporal signal, with held-out R^2 values of 0.71-0.87 for city coordinates and 0.48-0.52 for historical birth years. Semantic-neighbor analyses and targeted subspace ablations show that these signals depend strongly on interpretable lexical gradients, especially country names and climate-related vocabulary. These findings suggest that ordinary word co-occurrence preserves richer spatial, temporal, and environmental structure than is often assumed, revealing a remarkable and underappreciated capacity of simple static embeddings to preserve world-shaped structure from text alone. Linear probe recoverability alone therefore does not establish a representational move beyond text.

词嵌入语义梯度地理结构时间信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。