用大模型构建可解释的世界模型,零样本支持多智能体决策。
PIANIST: Learning Partially Observable World Models with LLMs for Multi-Agent Decision Making
- 将世界模型拆解为七类直观组件,实现零样本生成
- 仅凭自然语言描述即可生成可用于快速模拟的模型
- 适用于语言与非语言动作的复杂决策任务
大模型中有效提取用于复杂决策任务的世界知识仍具挑战。我们提出PIANIST框架,将世界模型分解为七个直观组件,有利于大模型零样本生成。仅需游戏的自然语言描述和观测输入格式,该方法即可生成可用于快速高效蒙特卡洛树搜索(MCTS)模拟的工作世界模型。实验表明,该方法在两款不同游戏中表现良好,均考验智能体的规划与决策能力,且支持基于语言与非语言的动作执行,无需领域特定训练数据或显式定义的世界模型。
原文摘要 · Abstract (English)
Effective extraction of the world knowledge in LLMs for complex decision-making tasks remains a challenge. We propose a framework PIANIST for decomposing the world model into seven intuitive components conducive to zero-shot LLM generation. Given only the natural language description of the game and how input observations are formatted, our method can generate a working world model for fast and efficient MCTS simulation. We show that our method works well on two different games that challenge the planning and decision making skills of the agent for both language and non-language based action taking, without any training on domain-specific training data or explicitly defined world model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。