长时序语言智能体的世界模型会突然崩溃,像水沸腾一样。
World-Model Collapse as a Phase Transition

- 通过精确状态任务研究世界模型随参数变化的相变行为
- 发现存在解的平台、窄过渡带和崩溃底层的相图结构
- 适合关注长期决策智能体鲁棒性与模型失效机制的研究者
水在加热过程中看似不变,达到临界点后突然沸腾。我们探讨长时序语言智能体的隐式世界模型是否也存在类似相变。在某些参数设置下,微小的状态负载变化或增加一步规划长度,行为几乎不变;而靠近临界边界时,相同变化却引发世界模型的突然崩溃。我们在一个具有精确每步真值状态的确定性任务族中研究该现象。通过大规模网格搜索状态基数、依赖密度、规划范围、分支度、观测模式和突变率,揭示出相图结构:解的平台区、狭窄的过渡带和崩溃底层。每步轨迹显示,世界状态保真度先于动作有效性失效,说明智能体并非选错动作,而是基于被污染的世界模型行动。更强模型可移动临界边界,但无法消除定性相变。这些结果使世界模型崩溃成为长时序智能体可测量的关键瓶颈。
原文摘要 · Abstract (English)
Water looks unchanged as it warms, then at a critical point it boils. We ask whether long-horizon language agents show an analogous transition in their implicit world models. In some parameter settings, changing state load by a small amount, or adding a single step of horizon, leaves behavior nearly unchanged; near a critical boundary, the same small change causes a sudden world collapse. We study this effect in a deterministic task family with exact per-step gold state. A large grid search over state cardinality, dependency density, horizon, branching, observation mode, and mutation rate reveals a phase diagram: a solved plateau, a narrow transition band, and a collapse floor. Per-step traces show the mechanism: world-state fidelity fails before action validity, so the agent is not merely choosing a bad action; it is acting from a corrupted world. Stronger models translate the critical boundary but do not remove the qualitative transition. These results make world-model collapse a measurable bottleneck for long-horizon agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。