评测智能体世界模型能否回答涉及全局环境的复杂问题。
Benchmarking World-Model Learning with Environment-Level Queries
- 设计新评估协议,测试模型对全局环境问题的回应能力。
- 人类在43个网格世界任务中显著优于现有模型。
- 适合研究具身智能与世界模型泛化能力的学者参考。
世界模型是实现灵活推理与规划的AI智能体的核心。然而当前评估仅关注可观测交互的属性(如帧预测或任务回报),未检验学习模型是否支持对环境的多样化查询。相比之下,人类构建的是通用模型,可回答涉及全局结构与反事实后果的问题。本文提出《WorldTest》:一种评估智能体是否具备支持多种环境级查询能力的协议——即答案依赖于整个环境而非仅观测轨迹的问题。这些查询可指向无法由单一轨迹分布决定的属性(如可达性或干预影响)。我们将其具体化为《AutumnBench》,包含43个交互式网格世界环境和129项任务,覆盖三类查询类型,适用于人类与学习智能体。517名人类参与者与五种前沿模型的实验表明,人类表现远超模型,差距归因于探索策略与信念更新机制的差异。AutumnBench为网格世界中的世界模型学习提供了评估框架,而WorldTest则为扩展至更复杂领域提供模板。
原文摘要 · Abstract (English)
World models are central to building AI agents capable of flexible reasoning and planning. Yet current evaluations (i) test only properties measurable from observed interactions, such as next-frame prediction or task return, and (ii) do not test whether a learned model supports diverse queries about the environment. In contrast, humans build $\textit{general-purpose}$ models that can answer many different questions about an environment$\unicode{x2014}$including questions that require understanding global structure and counterfactual consequences. We propose $\textit{WorldTest}$: a protocol for evaluating whether agents learn models that support multiple $\textit{environment-level queries}\unicode{x2014}$questions whose answers depend on properties of the full environment, not just observed trajectories. Individually, these queries can target properties (e.g., reachability or the effects of interventions) that no single rollout distribution determines. Collectively, they assess model generality across query types. We instantiate WorldTest as $\textit{AutumnBench}$, a benchmark of 43 interactive grid-world environments and 129 tasks across three query families for both humans and learning agents. Experiments with 517 human participants and five frontier models show that humans substantially outperform these models, a gap we attribute to differences in exploration and belief updating. AutumnBench provides a framework for evaluating world-model learning in grid-world environments with environment-level queries, and WorldTest provides a template for extending such evaluations to richer domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。