arXiv:2605.23238cs.AIcs.GT2026-05

用动态生成的博弈环境评估大模型战略推理能力,发现顶尖模型存在局部不稳定性。

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models

论文配图:GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models
图 1 · 摘自论文原文
  • 通过程序化生成多种扑克类博弈环境,实现持续可更新的评测
  • 9个主流大模型在超3.6万场对战中表现差异显著,新模型平均得分更高
  • 揭示模型在局部博弈中的波动性,为部署风险提供新诊断维度

大型语言模型(LLMs)正越来越多地被用作市场、拍卖和竞价场景中的经济主体。预测其在特定部署中的行为极具挑战性。现有战略推理基准仅测试固定经典博弈,易因性能饱和而失效,且难以推广至真实复杂多变的战略环境。本文提出GENSTRAT,采用程序化生成的战略环境来解决上述问题。具体而言,我们构建了一个由两玩家零和不完全信息扑克游戏组成的分布,可按需生成新游戏,确保评测永续且抗污染。我们结合能力剖面方法,从六个维度(状态空间、时间深度、信息敏感度、对手建模、风险偏好、脆弱性)分解模型能力。引入“锯齿度”指标衡量分布内平滑性,检测模型在策略相似游戏间优势突变现象。从2000个生成游戏中抽取50个作为基准,对九个前沿及开源大模型进行超过36,000场对战测试。结果显示,最新前沿模型平均得分更高;但得分相近的模型展现出截然不同的能力谱系,前两名(gpt-5和claude)在局部博弈中表现出明显波动性,而第三名(gemini-3.1-pro)则更稳定。能力剖面与锯齿度共同提供了超越总分排名的部署相关诊断。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed as economic agents in marketplaces, auctions, and bidding settings. Anticipating their behavior in any specific deployment is hard. Existing strategic-reasoning benchmarks evaluate models on fixed canonical games. These benchmarks may saturate as the frontier improves, and they do not allow evaluators to generalize with confidence from benchmark performance to the varied and messy strategic environments that actual deployments involve. We introduce GENSTRAT, which uses procedurally generated strategic environments to address these challenges. Concretely, we generate a distribution of two-player zero-sum imperfect-information card games. The generator can draw fresh games on demand, allowing for evergreen evaluation and resistance to contamination. We pair the game distribution with a capability-profile methodology that decomposes model competence across six axes (state space, temporal depth, information sensitivity, opponent modeling, risk, and brittleness). We also introduce a jaggedness measure of within-distribution smoothness that detects when a model's advantage jumps unpredictably between strategically similar games. We sample 50 benchmark games from a 2,000-game generated pool and evaluate nine frontier and open-weight LLMs in a head-to-head tournament with over 36,000 matches. Newer frontier-tier models score higher on average. Beyond that average, models with near-identical overall strength show qualitatively different capability profiles, and two of the top three leaderboard models (gpt-5 and claude) are noticeably more locally volatile than the third (gemini-3.1-pro), despite being close in overall strength. Together, the capability profile and the jaggedness measure give a deployment-relevant diagnostic that the overall ranking alone cannot provide.

战略推理博弈论大模型评估动态基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。