arXiv:2602.00528cs.AI2026-02中稿 · ICLR被引 8

LLM打牌不如专业算法,新框架让其推理更贴近博弈论。

How Far Are LLMs from Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use

  • 引入外部求解器生成符合博弈论的行动,提升推理一致性。
  • 实验显示新框架在德州扑克中表现超越现有方法,接近职业水平。
  • 适合研究大模型策略推理、博弈智能与工具增强方向的学者。

随着大型语言模型(LLMs)在高风险领域应用日益广泛,其在不确定环境下的战略推理能力变得至关重要。扑克提供了一个严格的测试平台,不仅要求精准的行动,还需遵循严谨的博弈论推理。本文系统评估了LLMs在多种真实扑克任务中的表现,涵盖游戏结果与推理过程。分析揭示:LLMs在对抗传统算法时表现不足,存在三大缺陷:依赖启发式规则、事实理解错误,以及‘知行不一’——行动与推理脱节。初步尝试通过行为克隆与步骤级强化学习改善推理风格,但难以实现准确的博弈论决策。为此,我们提出ToolPoker框架,结合外部求解器生成符合博弈论最优(GTO)的行动,并辅以更精确的职业风格解释。实验表明,ToolPoker在游戏表现上达到当前最佳水平,且推理轨迹高度契合博弈论原则。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) are increasingly applied in high-stakes domains, their ability to reason strategically under uncertainty becomes critical. Poker provides a rigorous testbed, requiring not only strong actions but also principled, game-theoretic reasoning. In this paper, we conduct a systematic study of LLMs in multiple realistic poker tasks, evaluating both gameplay outcomes and reasoning traces. Our analysis reveals LLMs fail to compete against traditional algorithms and identifies three recurring flaws: reliance on heuristics, factual misunderstandings, and a "knowing-doing" gap where actions diverge from reasoning. An initial attempt with behavior cloning and step-level reinforcement learning improves reasoning style but remains insufficient for accurate game-theoretic play. Motivated by these limitations, we propose ToolPoker, a tool-integrated reasoning framework that combines external solvers for GTO-consistent actions with more precise professional-style explanations. Experiments demonstrate that ToolPoker achieves state-of-the-art gameplay while producing reasoning traces that closely reflect game-theoretic principles.

博弈论大模型推理工具增强扑克智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。