构建统一游戏评测平台,评估视觉语言模型在多类游戏中的持续进化能力。
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics

- 基于UE5打造12款新游戏,支持单人、对战和合作模式,统一动作接口。
- 引入改进动态曲线,让模型自我反思并逐步优化技能提示,提升表现。
- 不仅看初始成绩,还追踪进化过程与泛化能力,适合评估真实智能体。
视觉语言模型(VLM)代理正被越来越多地部署于交互式游戏环境。然而,现有针对VLM代理的游戏基准测试通常仅报告每个(代理,游戏)组合的首次尝试得分,聚焦于单代理独奏模式,并缺乏对异构代理类别(商业VLM、开源权重VLM、专用游戏策略)在统一标准下评估的协议。为此,我们提出OmniGameArena,一个实时基准测试平台,包含12个新构建的虚幻引擎5(Unreal Engine 5)游戏,涵盖独奏(7个)、玩家对战(3个)和合作(2个),具备统一的动作接口。同时引入改进动态曲线(IDC),一种代理反思机制,其中使用工具的反思型大模型可自主在多轮中优化有限制的技能提示。除了冷启动排行榜分数外,IDC为每个(代理,游戏)组合揭示两个额外可观测指标:得分随反思轮次的演变过程,以及所学技能在未见任务变体上的表现。我们报告了12个VLM代理在冷启动排行榜上的结果,以及4个顶级代理在IDC下的表现。
原文摘要 · Abstract (English)
Vision-language model (VLM) agents are increasingly deployed in interactive game environments. Yet game benchmarks for VLM agents typically report a single first-attempt score per (agent, game) pair, focus on single-agent Solo play, and lack unified protocols for evaluating heterogeneous agent classes (commercial VLMs, open-weight VLMs, and specialized game policies) on the same footing. We address these gaps with OmniGameArena, a real-time benchmark of twelve newly built Unreal Engine 5 games spanning Solo (7), PvP (3), and Coop (2) with unified action interfaces, and the Improvement Dynamics Curve (IDC), an agentic-reflection harness in which a tool-using reflector LLM autonomously refines a bounded skill prompt across multiple rounds. Beyond cold-start leaderboard scores, IDC exposes two additional observables for each (agent, game) pair: how the score evolves across reflection rounds, and how the learned skill behaves on held-out task variants. We report these observables for twelve VLM agents on the cold-start leaderboard and four top agents under IDC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。