arXiv:2506.10264cs.AI2025-06被引 5

用兵棋推演评估大模型的战略推理能力,填补智能体博弈短板。

WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models

  • 以兵棋为场景构建战略推理评测基准,覆盖环境感知、对手建模与策略生成。
  • 提出S-POE架构,系统测试模型在动态对抗中的决策与反事实推理能力。
  • 适合研究大模型战略智能、多智能体博弈与决策生成的学者使用。

大型语言模型(LLMs)在数学、符号与常识推理任务上取得显著进展,但在高级人类认知的核心能力——战略推理方面仍缺乏系统评估。本文提出WGSR-Bench,首个基于兵棋场景的博弈论战略推理评测基准。兵棋融合环境不确定性、对抗动态与非唯一策略选择,是评估多智能体决策、意图推断与反事实推理的理想场景。该基准围绕三大核心任务设计:环境态势感知、对手风险建模与策略生成,构成S-POE架构,系统评估战略推理能力。最后,构建基于LLM的兵棋智能体,整合各模块实现综合评估。本工作旨在揭示当前先进大模型在博弈论战略推理中的优劣,并推动大模型驱动的战略智能研究。

原文摘要 · Abstract (English)

Recent breakthroughs in Large Language Models (LLMs) have led to a qualitative leap in artificial intelligence' s performance on reasoning tasks, particularly demonstrating remarkable capabilities in mathematical, symbolic, and commonsense reasoning. However, as a critical component of advanced human cognition, strategic reasoning, i.e., the ability to assess multi-agent behaviors in dynamic environments, formulate action plans, and adapt strategies, has yet to be systematically evaluated or modeled. To address this gap, this paper introduces WGSR-Bench, the first strategy reasoning benchmark for LLMs using wargame as its evaluation environment. Wargame, a quintessential high-complexity strategic scenario, integrates environmental uncertainty, adversarial dynamics, and non-unique strategic choices, making it an effective testbed for assessing LLMs' capabilities in multi-agent decision-making, intent inference, and counterfactual reasoning. WGSR-Bench designs test samples around three core tasks, i.e., Environmental situation awareness, Opponent risk modeling and Policy generation, which serve as the core S-POE architecture, to systematically assess main abilities of strategic reasoning. Finally, an LLM-based wargame agent is designed to integrate these parts for a comprehensive strategy reasoning assessment. With WGSR-Bench, we hope to assess the strengths and limitations of state-of-the-art LLMs in game-theoretic strategic reasoning and to advance research in large model-driven strategic intelligence.

战略推理兵棋推演多智能体大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。