arXiv:2606.13815cs.AIcs.CL2026-06

用九维认知评估框架,揭示大模型在德州扑克中的真实推理能力

Poker Arena: Multi-Axis Profiling of Strategic Reasoning and Memory in LLMs

论文配图:Poker Arena: Multi-Axis Profiling of Strategic Reasoning and Memory in LLMs
图 1 · 摘自论文原文
  • 构建三层次记忆架构与九维度策略评估体系
  • Claude Opus 4.6赢15730筹码但综合评分仅第五
  • 跨维度一致性比单轴峰值更重要

战略推理在不确定性下的决策关乎谈判、金融与政策,但现有博弈评测将多元推理能力简化为单一数值,掩盖了前沿大模型的真实能力结构。我们提出Poker Arena——一个无限制德州扑克赛事平台,结合三层记忆架构(手内、会话、跨会话)与九轴认知评估体系,分解出下注尺度校准、位置意识等可解释的推理维度。在7个前沿模型上进行50场、每场1000手的测试,并控制记忆模块消融;筹码收益与综合轴得分排序不一致:Claude Opus 4.6以+15,730筹码夺冠并获14次第一,但在平均轴得分中仅列第五;持久记忆对部分模型有益,对另一些则有害。结果表明,多轴评估能揭示单维度排行榜无法捕捉的能力结构,跨维度一致性优于单一维度的峰值表现。

原文摘要 · Abstract (English)

Strategic reasoning under uncertainty underpins consequential decisions in negotiation, finance, and policy, but prevailing game-play benchmarks collapse heterogeneous reasoning dimensions into a single scalar, leaving the capability structure of frontier LLMs unexamined. We introduce Poker Arena, a no-limit Texas Hold'em tournament platform that couples a three-layer memory architecture (within-hand, session, and cross-session) with a nine-axis cognitive profile decomposing strategic reasoning into interpretable dimensions such as bet-sizing calibration and positional awareness. We evaluate seven frontier models across 50 sessions of 1,000 hands and a controlled memory ablation; tournament chips and aggregate axis score order the field differently: Claude Opus 4.6 wins +$15,730 chips with 14 first-place finishes, yet ranks only fifth of seven on mean axis score, while persistent memory helps some models and hurts others. These findings show that multi-axis evaluation surfaces capability structure that scalar leaderboards systematically misrank, with cross-dimensional consistency outweighing peak performance on any single axis.

策略推理认知评估大模型测评扑克博弈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。