用生态演化视角分析语言模型如何捕捉世界状态,揭示其表征选择机制。
Task Ecologies and the Evolution of World-Tracking Representations in Large Language Models

- 基于贝叶斯最优性推导出语言模型表征的生态真实性标准
- 发现训练生态中等价类的商划分是零误差的最小复杂度解
- 提出小模型可作为实验工具,研究表征选择的演化规律
我们把语言模型视为演化的模型生物,探究自回归下一个词学习何时会选择世界追踪表征。对于任何潜在世界状态的编码,贝叶斯最优的下一个词交叉熵可分解为不可约条件熵与一个Jensen-Shannon过剩项。该过剩项仅在编码保留训练生态等价类时消失。这给出了语言模型生态真实性的精确定义,并识别出最小复杂度零过剩解为训练等价类的商划分。随后我们确定该固定编码分析适用于变换器家族:冻结的密集型和冻结的专家混合模型满足此条件,上下文学习不会扩大模型的分离集,而任务适配会破坏前提。框架预测两种典型失败模式:简单性压力优先消除低收益区分,且训练最优模型在更精细的部署生态中仍可能产生正过剩。通过条件动态扩展表明,在显式遗传、变异和选择假设下,模型间选择与后训练可恢复这些差距区分。精确有限生态检验和受控microGPT实验验证了静态分解、分裂-合并阈值、离生态失败模式及双生态救援机制,在相关量可直接观测的范围内成立。目标并非大规模建模前沿系统,而是利用小型语言模型作为理论研究表征选择的实验室生物。
原文摘要 · Abstract (English)
We study language models as evolving model organisms and ask when autoregressive next-token learning selects for world-tracking representations. For any encoding of latent world states, the Bayes-optimal next-token cross-entropy decomposes into the irreducible conditional entropy plus a Jensen--Shannon excess term. That excess vanishes if and only if the encoding preserves the training ecology's equivalence classes. This yields a precise notion of ecological veridicality for language models and identifies the minimum-complexity zero-excess solution as the quotient partition by training equivalence. We then determine when this fixed-encoding analysis applies to transformer families: frozen dense and frozen Mixture-of-Experts transformers satisfy it, in-context learning does not enlarge the model's separation set, and per-task adaptation breaks the premise. The framework predicts two characteristic failure modes: simplicity pressure preferentially removes low-gain distinctions, and training-optimal models can still incur positive excess on deployment ecologies that refine the training ecology. A conditional dynamic extension shows how inter-model selection and post-training can recover such gap distinctions under explicit heredity, variation, and selection assumptions. Exact finite-ecology checks and controlled microgpt experiments validate the static decomposition, split-merge threshold, off-ecology failure pattern, and two-ecology rescue mechanism in a regime where the relevant quantities are directly observable. The goal is not to model frontier systems at scale, but to use small language models as laboratory organisms for theory about representational selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。