arXiv:2507.22281cs.AIcs.CL2025-07EMNLP被引 5

让大模型世界模型与探索过程同步进化,提升复杂任务规划能力

CoEx -- Co-evolving World-model and Exploration

  • 用分层状态抽象实现大模型推理与动态世界模型协同演化
  • 在ALFWorld等环境上显著优于现有方法,提升规划与探索效果
  • 适合研究智能体自主学习、复杂环境交互的学者参考

现代大模型智能体的规划依赖于预训练中获得的静态内部世界模型。然而,现有设计难以将新观测有效融入世界模型的动态更新,导致模型逐渐偏离真实世界状态,生成错误计划。我们提出一种分层智能体架构CoEx,通过分层状态抽象使大模型规划与动态世界模型共同演化。CoEx利用大模型推理协调由子目标构成的动态计划,并通过持续学习将子目标经验以神经符号信念状态(包含文本推断与代码化符号记忆)形式融入持久的世界模型。我们在包含丰富环境和复杂任务的ALFWorld、PDDL和Jericho等多个场景中评估该智能体,实验表明CoEx在规划与探索方面均优于现有范式。

原文摘要 · Abstract (English)

Planning in modern LLM agents relies on the utilization of LLM as an internal world model, acquired during pretraining. However, existing agent designs fail to effectively assimilate new observations into dynamic updates of the world model. This reliance on the LLM's static internal world model is progressively prone to misalignment with the underlying true state of the world, leading to the generation of divergent and erroneous plans. We introduce a hierarchical agent architecture, CoEx, in which hierarchical state abstraction allows LLM planning to co-evolve with a dynamically updated model of the world. CoEx plans and interacts with the world by using LLM reasoning to orchestrate dynamic plans consisting of subgoals, and its learning mechanism continuously incorporates these subgoal experiences into a persistent world model in the form of a neurosymbolic belief state, comprising textual inferences and code-based symbolic memory. We evaluate our agent across a diverse set of agent scenarios involving rich environments and complex tasks including ALFWorld, PDDL, and Jericho. Our experiments show that CoEx outperforms existing agent paradigms in planning and exploration.

智能体世界模型规划大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。