用大模型做世界模型,能独立决策,但长期任务表现差且不稳定。
LLM-Based World Models Can Make Decisions Solely, But Rigorous Evaluations are Needed
- 构建31个环境+规则策略,从决策角度评估大模型世界模型
- GPT-4o在需领域知识的任务中远超GPT-4o-mini,长时决策性能下降
- 世界模型多功能组合会引入性能不稳,需严谨评估
世界模型是决策的关键模块,MuZero和Dreamer在复杂任务中取得显著成功。近期工作利用大语言模型(LLMs)作为通用世界模拟器,因其泛化能力,也用于推理规划(RAP)和思维树(ToT)中的反思性推理。然而,现有评估仅将其视为通用模拟器或辅助规划的模块。本文从决策视角提出全面评估框架:采用Wang et al. (2023; 2024)的31个多样化环境,并为每个环境构建规则策略进行评估。设计三项核心任务——策略验证、动作提议和策略规划,使世界模型可独立完成决策。对先进LLMs(GPT-4o与GPT-4o-mini)在不同设置下执行三类任务进行综合评估。关键发现包括:i) GPT-4o在需领域知识的任务中显著优于GPT-4o-mini;ii) 长期决策任务中性能明显下降;iii) 多功能组合导致性能额外不稳定。
原文摘要 · Abstract (English)
World model emerges as a key module in decision making, where MuZero and Dreamer achieve remarkable successes in complex tasks. Recent work leverages Large Language Models (LLMs) as general world simulators to simulate the dynamics of the world due to their generalizability. LLMs also serve as the world model for deliberative reasoning in Reasoning via Planning (RAP) and Tree of Thought (ToT). However, the world models are either evaluated as a general world simulator, or as a functional module of the agent, i.e., predicting the transitions to assist the planning. In this work, we propose a comprehensive evaluation of the world models with LLMs from the decision making perspective. Specifically, we leverage the 31 diverse environments from (Wang et al., 2023;2024) and curate the rule-based policy of each environment for the diverse evaluation. Then, we design three main tasks, i.e., policy verification, action proposal, and policy planning, where the world models can be used for decision making solely. Finally, we conduct the comprehensive evaluation of the advanced LLMs, i.e., GPT-4o and GPT-4o-mini, on the environments for the three main tasks under various settings. The key observations include: i) GPT-4o significantly outperforms GPT-4o-mini on the three main tasks, especially for the tasks which require the domain knowledge, ii) the performance of the world model with LLM will be decreased for long-term decision-making tasks, and iii) the combination of different functionalities of the world model will brings additional unstabilities of the performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。