用兵棋推演评估大模型的战略推理能力,填补智能体博弈短板。
WGSR-Bench: Wargame-based Game-theoretic Strategic Reasoning Benchmark for Large Language Models
- 以兵棋为场景构建战略推理评测基准,覆盖环境感知、对手建模与策略生成。
- 提出S-POE架构,系统测试模型在动态对抗中的决策与反事实推理能力。
- 适合研究大模型战略智能、多智能体博弈与决策生成的学者使用。
大型语言模型(LLMs)在数学、符号与常识推理任务上取得显著进展,但在高级人类认知的核心能力——战略推理方面仍缺乏系统评估。本文提出WGSR-Bench,首个基于兵棋场景的博弈论战略推理评测基准。兵棋融合环境不确定性、对抗动态与非唯一策略选择,是评估多智能体决策、意图推断与反事实推理的理想场景。该基准围绕三大核心任务设计:环境态势感知、对手风险建模与策略生成,构成S-POE架构,系统评估战略推理能力。最后,构建基于LLM的兵棋智能体,整合各模块实现综合评估。本工作旨在揭示当前先进大模型在博弈论战略推理中的优劣,并推动大模型驱动的战略智能研究。
原文摘要 · Abstract (English)
Recent breakthroughs in Large Language Models (LLMs) have led to a qualitative leap in artificial intelligence' s performance on reasoning tasks, particularly demonstrating remarkable capabilities in mathematical, symbolic, and commonsense reasoning. However, as a critical component of advanced human cognition, strategic reasoning, i.e., the ability to assess multi-agent behaviors in dynamic environments, formulate action plans, and adapt strategies, has yet to be systematically evaluated or modeled. To address this gap, this paper introduces WGSR-Bench, the first strategy reasoning benchmark for LLMs using wargame as its evaluation environment. Wargame, a quintessential high-complexity strategic scenario, integrates environmental uncertainty, adversarial dynamics, and non-unique strategic choices, making it an effective testbed for assessing LLMs' capabilities in multi-agent decision-making, intent inference, and counterfactual reasoning. WGSR-Bench designs test samples around three core tasks, i.e., Environmental situation awareness, Opponent risk modeling and Policy generation, which serve as the core S-POE architecture, to systematically assess main abilities of strategic reasoning. Finally, an LLM-based wargame agent is designed to integrate these parts for a comprehensive strategy reasoning assessment. With WGSR-Bench, we hope to assess the strengths and limitations of state-of-the-art LLMs in game-theoretic strategic reasoning and to advance research in large model-driven strategic intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。