让大模型学会提前模拟未来,做出更聪明的决策。
Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning

- 用三阶段训练让模型自动生成未来状态和成功预测
- 在搜索与数学任务中表现优于其他方法
- 适合需要长期规划的智能体研究者
大型语言模型代理在序列决策中表现出色,但在长周期任务中仍具反应性。人类能通过‘如果……会怎样’的推理提前评估计划,而标准代理缺乏内部世界模型来模拟未来结果。为此,我们提出通过训练单一自回归模型,生成前瞻性状态演进和计划条件的成功估计——一种文本形式的Q值。关键发现是存在格式-能力鸿沟:仅在后训练阶段对前瞻轨迹进行微调,只会导致表面模仿,缺乏真实预测基础。为此,我们引入三阶段训练范式:(i) 世界模型代理中段训练(WM-AMT),注入潜在预测能力;(ii) 格式诱发SFT(FE-SFT),结构化该能力;(iii) 未来条件强化学习(FC-RL),优化生成模拟的校准性与实用性。在搜索与数学推理任务上的评估显示,本方法持续优于其他训练基线。结果表明,有效的内部世界建模需以能力优先的训练流程,才能实现有根基且校准良好的前瞻性判断。
原文摘要 · Abstract (English)
Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remains fundamentally reactive in long-horizon tasks. Unlike humans who employ "what-if" reasoning to evaluate potential plans before commitment, standard agents lack an internal world model to simulate future outcomes. Therefore, we propose to internalize future-aware planning by training a single autoregressive model to verbalize both a prospective state rollout and a plan-conditioned success estimate-a textual analogue of the Q-value. Crucially, we identify a format-capability gap: simply fine-tuning agents on look-ahead traces during post-training leads to superficial mimicry of foresight without genuine predictive grounding. To bridge this gap, we introduce a three-stage training paradigm: (i) World Model Agentic Mid-Training (WM-AMT) to inject latent predictive capabilities into the policy; (ii) Format-Eliciting SFT (FE-SFT) to structure this injected capability; and (iii) Foresight-Conditioned Reinforcement Learning (FC-RL) to refine the calibration and utility of the generated simulations. Evaluated on search and mathematical reasoning tasks, our approach consistently outperforms other training baselines. Our results demonstrate that effective internal world modeling in LLM agents requires a capability-first training pipeline to achieve grounded and calibrated foresight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。