构建标准化游戏评测框架,验证多模态大模型在游戏中的综合能力
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
- 设计双接口游戏代理:直接操控与语义动作解析
- 覆盖34款游戏170项任务,结果均未达人类水平
- 支持可重复验证,适合评估游戏代理的鲁棒性与泛化性
为实现真实世界交互的具身通用智能体,多模态大语言模型(MLLM)代理仍面临延迟高、反馈稀疏和不可逆错误等挑战。视频游戏提供丰富的视觉观察与闭环交互环境,对精细感知、长时程规划和精准控制提出要求。然而,当前系统性评估受限于动作接口异构与启发式验证方式。为此,我们提出GameWorld,一个面向浏览器环境中多模态大模型通用游戏代理的标准化、可验证评估基准。研究两种代理接口:(i) 直接发出键盘鼠标指令的计算机使用代理;(ii) 通过确定性语义动作解析在语义空间行动的通用多模态代理。GameWorld包含34种多样游戏和170项任务,每项任务配备状态可验证指标,实现基于结果的评估。18组模型-接口组合的结果表明,即使表现最佳的代理也远未达到人类水平。多次全基准重跑实验验证了基准的鲁棒性,进一步对实时交互、上下文记忆敏感性和动作有效性分析揭示了更多挑战。GameWorld通过提供标准化、可验证、可复现的评估框架,为推进多模态游戏代理研究奠定坚实基础。
原文摘要 · Abstract (English)
Towards an embodied generalist for real-world interaction, Multimodal Large Language Model (MLLM) agents still suffer from challenging latency, sparse feedback, and irreversible mistakes. Video games offer an ideal testbed with rich visual observations and closed-loop interaction, demanding fine-grained perception, long-horizon planning, and precise control. However, systematically evaluating these capabilities is currently hindered by heterogeneous action interfaces and heuristic verification. To this end, we introduce GameWorld, a benchmark designed for standardized and verifiable evaluation of MLLMs as generalist game agents in browser environments. Two game agent interfaces are studied: (i) computer-use agents that directly emit keyboard and mouse controls, and (ii) generalist multimodal agents that act in a semantic action space via deterministic Semantic Action Parsing. GameWorld contains 34 diverse games and 170 tasks, each paired with state-verifiable metrics for outcome-based evaluation. The results across 18 model-interface pairs suggest that even the best performing agent is far from achieving human capabilities on video games. Extensive experiments of repeated full-benchmark reruns demonstrate the robustness of the benchmark, while further studies on real-time interaction, context-memory sensitivity, and action validity expose more challenges ahead for game agents. Together, by offering a standardized, verifiable, and reproducible evaluation framework, GameWorld lays a robust foundation for advancing research on multimodal game agents and beyond. The project page is at https://gameworld-bench.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。