arXiv:2603.23660cs.AI2026-03被引 4

打造首个公开的德州扑克基准测试,评估AI在不完全信息下的决策能力。

GTO Wizard Benchmark

  • 用GTO Wizard AI作为强基准,量化评估博弈智能体性能。
  • 引入AIVAT技术,使评估效率提升10倍,显著降低方差影响。
  • 首次评测大模型零样本推理能力,揭示其在隐藏状态推理上的短板。

我们提出GTO Wizard Benchmark,一个公开的API与标准化评估框架,用于评测两人无限制德州扑克(HUNL)中的算法性能。该基准以最先进的超人类扑克代理GTO Wizard AI为对手,其击败2018年计算机扑克大赛冠军Slumbot,平均胜率达19.4 ± 4.1比分/百手。由于扑克评估中方差问题突出,我们集成AIVAT这一可证明无偏的方差缩减技术,使相同统计显著性所需对局数减少至传统蒙特卡洛方法的十分之一。我们对当前顶尖大语言模型(包括GPT-5.4、Claude Opus 4.6、Gemini 3.1 Pro、Grok 4等)进行零样本条件下全面评测。初步结果显示大模型推理能力近年有显著进步,但所有模型仍远低于本基准设定的基线水平。定性分析表明,模型在状态表示与隐状态推理方面存在明显改进空间。该基准为研究部分可观测多智能体系统中的规划与推理提供了精确可量化的评估环境。

原文摘要 · Abstract (English)

We introduce GTO Wizard Benchmark, a public API and standardized evaluation framework for benchmarking algorithms in Heads-Up No-Limit Texas Hold'em (HUNL). The benchmark evaluates agents against GTO Wizard AI, a state-of-the-art superhuman poker agent that approximates Nash Equilibria, and defeated Slumbot, the 2018 Annual Computer Poker Competition champion and previous strongest publicly accessible HUNL benchmark, by $19.4$ $\pm$ $4.1$ bb/100. Variance is a fundamental challenge in poker evaluation; we address this by integrating AIVAT, a provably unbiased variance reduction technique that achieves equivalent statistical significance with ten times fewer hands than naive Monte Carlo evaluation. We conduct a comprehensive benchmarking study of state-of-the-art large language models under zero-shot conditions, including GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro, Grok 4, and others. Initial results and analysis reveal dramatic progress in LLM reasoning over recent years, yet all models remain far below the baseline established by our benchmark. Qualitative analysis reveals clear opportunities for improvement, including representation and the ability to reason over hidden states. This benchmark provides researchers with a precise and quantifiable setting to evaluate advances in planning and reasoning in multi-agent systems with partial observability.

扑克博弈大模型评估零样本推理多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。