用RTS游戏测试视觉语言模型的战略推理能力,发现其在复杂协作中表现不佳。
RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models

- 基于大规模RTS游戏构建动态评估框架,支持多样化对战和诊断性小关卡。
- 模型在高协同需求和任务规模增大时性能显著下降,暴露战略推理短板。
- 支持自演化生成新关卡,适合研究多智能体协作与长程规划的学者使用。
现代视觉语言模型(VLMs)在竞争与合作场景中常难以进行战略推理,即在不确定性下预判并影响其他智能体行为。实时策略(RTS)游戏是检验此能力的理想场景,因其要求与队友协作、适应对手策略,并在部分可观测条件下进行长期规划。然而,现有RTS基准评估范围有限,缺乏系统性能力诊断,且场景固定。为此,我们提出RTSGameBench,基于大规模RTS游戏Beyond All Reason,其战场更大,策略多样性更高。该基准通过多样对战结构进行评估,利用针对单一战略能力设计的小关卡实现诊断性测评,并引入自演化生成框架,将自由文本查询转化为新小关卡,实现持续扩展。此外,为使VLM在大规模RTS中运行,我们提供RTSGameAgent,采用状态机加代理记忆管理单位行为。实验验证,多个前沿VLM在需紧密协作、多智能体协调及任务规模增大时表现不佳。
原文摘要 · Abstract (English)
Modern Vision-Language Models (VLMs) often struggle with strategic reasoning, i.e., anticipating and influencing other agents' actions, under uncertainty in competitive and cooperative settings. Real-time strategy (RTS) games can be a natural testbed for diagnosing this limitation, as they demand coordination with allies, adaptation to opponents' strategy, and long-horizon planning under partial observability. However, existing RTS benchmarks offer limited evaluation scope, lack systematic competency diagnosis, and remain fixed in the pre-designed scenario coverage. To address these limitations, we present RTSGameBench, which is built on Beyond All Reason, a large-scale RTS game with an expanded battlefield that demands broader strategy diversity than the existing testbeds. The proposed benchmark provides evaluations through diverse gameplay across various matchup structures, diagnostic assessment via mini-games, each targeting an individual strategic competency, and extensible coverage via a self-evolving generation framework that converts free-form queries into new mini-games, improving over successive cycles. Additionally, for VLMs to operate in large-scale RTS games, we provide RTSGameAgent that manages units by an FSM with agentic memory. We empirically validate that multiple state-of-the-art VLMs do not perform well when matchups demand tighter coordination, multiagent coordination and when task scale increases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。