arXiv:2509.16610cs.CL2025-09EMNLP被引 11

用博弈论游戏测试大模型的战略思维能力,发现不同模型的决策差异。

LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts

  • 基于博弈论设计多智能体游戏环境,评估模型策略与互动行为。
  • 评测15个主流大模型,得分反映其推理与战略能力强弱。
  • 适合关注模型社交推理与复杂决策能力的研究者使用。

随着大语言模型在各类任务中的持续进步,单一指标已难以全面评估其能力。为真正衡量大模型的智能水平,必须考察其交互动态与战略行为。我们提出 LLMsPark,一个基于博弈论的评估平台,用于测量大模型在经典博弈论场景下的决策策略与社会行为,提供多智能体环境以探索战略深度。该系统对15个领先的大模型(含商用与开源)进行交叉评估,采用排行榜与评分机制。得分越高,表明推理与战略能力越强,揭示出各模型间显著的行为模式与性能差异。本工作为评估大模型的战略智能提供了新视角,丰富了现有基准体系,并拓展了其在交互式、博弈论场景下的评估范畴。基准与排名已公开发布于 https://llmsparks.github.io/。

原文摘要 · Abstract (English)

As large language models (LLMs) advance across diverse tasks, the need for comprehensive evaluation beyond single metrics becomes increasingly important. To fully assess LLM intelligence, it is crucial to examine their interactive dynamics and strategic behaviors. We present LLMsPark, a game theory-based evaluation platform that measures LLMs' decision-making strategies and social behaviors in classic game-theoretic settings, providing a multi-agent environment to explore strategic depth. Our system cross-evaluates 15 leading LLMs (both commercial and open-source) using leaderboard rankings and scoring mechanisms. Higher scores reflect stronger reasoning and strategic capabilities, revealing distinct behavioral patterns and performance differences across models. This work introduces a novel perspective for evaluating LLMs' strategic intelligence, enriching existing benchmarks and broadening their assessment in interactive, game-theoretic scenarios. The benchmark and rankings are publicly available at https://llmsparks.github.io/.

大模型评估博弈论战略智能多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。