LLM Agent 的世界模型与策略通过互补的几何方式协同学习。
How do World Models and Policies Compose in LLM Agents? A Joint Spectral and Behavioral Account

- 通过参数更新的几何分解,揭示世界模型与策略的协同机制。
- 顺序训练的代理探索更广状态空间,且对输入扰动更鲁棒。
- 建议在后训练阶段设计两者间的接口以增强知识保留。
通过结合世界模型训练(下一状态预测)和策略训练(奖励最大化)的受控实验,我们探究了大语言模型智能体如何理解所处环境并掌握任务。通过分析模型的加性参数更新,发现有效世界模型更新为低秩,其输入特征子空间与策略更新共享,但输出方向几乎正交,无论单独或顺序训练。然而,在投影干预中,顺序更新在移除世界模型主导输入方向时表现出更强鲁棒性,表明其学习了替代输入路径。行为上,顺序训练的代理探索更广泛的状态与动作空间。进一步追问:策略训练是否充分保留世界知识?我们提出无需训练的融合方法,基于几何启发的输入基底与在线世界模型损失,结果优于未处理基线。研究提示,世界知识与任务能力可呈几何互补形式学习,未来后训练流程应优化二者接口。
原文摘要 · Abstract (English)
How do LLM agents come to both understand environments they act in and master tasks set within them? Through controlled experiments combining world-model training (next-state prediction) and policy training (reward maximization), we investigate this question. We dissect the resulting models through their additive parameter updates. Geometrically, we find effective world-model updates are low-rank and share an input-feature subspace with policy updates while writing to nearly orthogonal output directions, whether trained separately or sequentially. However, we find that, in projection interventions, the sequential update induces more robustness than separate policy RL when removing the world model's leading input directions, suggesting that it has learned alternative input pathways. Behaviorally, we find the sequentially trained agent explores a wider range of states and actions. Based on this, we ask: does policy training preserve world knowledge as well as it could? We probe this with training-free merging built on the geometrically motivated input basis plus an online world-model loss during policy RL, and show both improve over the untreated baseline. Our findings suggest world knowledge and task-directed ability can be learned in geometrically complementary forms, and that future post-training pipelines should consider how best to engineer the interface between them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。