测试AI能否像人一样用抽象动作规划长期任务。
WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
- 设计视觉基准,评估模型对抽象动作序列的推理能力。
- 当前顶尖模型在任务中准确率仅57%和38%,远低于人类完美表现。
- 引入动作等价物防作弊,确保评估更公平可靠。
人类拥有内在的‘世界模型’,可基于环境状态进行行动规划。人工智能代理同样需要此类世界模型以实现行动规划。然而,当前生成模型如何学习世界模型并在多样化环境中进行过程性规划仍不明确。为此,我们提出WorldPrediction,一个基于视频的基准,用于评估不同AI模型的世界建模与过程性规划能力。不同于以往侧重低层建模和机器人运动规划的基准,WorldPrediction首次强调具有时间与语义抽象的动作。给定初始与最终世界状态,任务是区分正确的单一动作(WorldPrediction-WM)或正确排序的动作序列(WorldPrediction-PP),从一系列反事实干扰项中选出。这种判别式任务设计使我们能够评估多种世界模型与规划器,并实现跨假设的全面比较。基准通过视觉观测表示状态与动作。为防止模型利用背景场景的低层连续性线索,我们提供‘动作等价物’——在不同情境下观察到的相同动作——作为候选选择。该基准基于部分可观测半马尔可夫决策过程(POMDP)的正式框架,确保评估的可靠性与鲁棒性。我们进行了广泛的人工筛选与验证,结果显示当前前沿模型在WorldPrediction-WM上准确率仅为57%,在WorldPrediction-PP上为38%,而人类可完美解决两项任务。
原文摘要 · Abstract (English)
Humans are known to have an internal "world model" that enables us to carry out action planning based on world states. AI agents need to have such a world model for action planning as well. It is not clear how current AI models, especially generative models, are able to learn such world models and carry out procedural planning in diverse environments. We introduce WorldPrediction, a video-based benchmark for evaluating world modeling and procedural planning capabilities of different AI models. In contrast to prior benchmarks that focus primarily on low-level world modeling and robotic motion planning, WorldPrediction is the first benchmark that emphasizes actions with temporal and semantic abstraction. Given initial and final world states, the task is to distinguish the proper action (WorldPrediction-WM) or the properly ordered sequence of actions (WorldPrediction-PP) from a set of counterfactual distractors. This discriminative task setup enable us to evaluate different types of world models and planners and realize a thorough comparison across different hypothesis. The benchmark represents states and actions using visual observations. In order to prevent models from exploiting low-level continuity cues in background scenes, we provide "action equivalents" - identical actions observed in different contexts - as candidates for selection. This benchmark is grounded in a formal framework of partially observable semi-MDP, ensuring better reliability and robustness of the evaluation. We conduct extensive human filtering and validation on our benchmark and show that current frontier models barely achieve 57% accuracy on WorldPrediction-WM and 38% on WorldPrediction-PP whereas humans are able to solve both tasks perfectly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。