arXiv:2501.08328cs.CLcs.AI2025-01AAAI被引 21

用扑克牌测试大模型策略能力,发现现役模型需微调才够专业。

PokerBench: Training Large Language Models to become Professional Poker Players

  • 构建1.1万种关键牌局场景,分翻前翻后评估模型表现。
  • 微调后模型胜率显著提升,分数与实战赢率正相关。
  • 揭示监督微调局限性,推动更优训练方法研究。

我们提出PokerBench——一个评估大语言模型(LLM)扑克博弈能力的基准。尽管大模型在传统自然语言任务中表现优异,但将其应用于扑克这类不完全信息博弈仍具挑战。扑克要求数学计算、推理规划、策略制定、博弈论理解及对人类心理的把握,是检验大模型复杂决策能力的理想场景。PokerBench由11,000个重要牌局构成,涵盖翻前与翻后阶段,与专业扑克选手合作开发。我们评估了GPT-4、ChatGPT 3.5及多个Llama和Gemma系列模型,发现所有主流模型在最优策略上均表现不足。经微调后,模型性能显著改善。通过让不同得分模型相互对战,验证了其得分与实际胜率的正相关性。对比微调模型与GPT-4的对局,揭示了简单监督微调在学习最优策略上的局限,表明需要更先进的训练方法。PokerBench为快速可靠评估大模型扑克能力提供了独特基准,并为研究大模型在复杂博弈中的进展提供了全面平台。

原文摘要 · Abstract (English)

We introduce PokerBench - a benchmark for evaluating the poker-playing abilities of large language models (LLMs). As LLMs excel in traditional NLP tasks, their application to complex, strategic games like poker poses a new challenge. Poker, an incomplete information game, demands a multitude of skills such as mathematics, reasoning, planning, strategy, and a deep understanding of game theory and human psychology. This makes Poker the ideal next frontier for large language models. PokerBench consists of a comprehensive compilation of 11,000 most important scenarios, split between pre-flop and post-flop play, developed in collaboration with trained poker players. We evaluate prominent models including GPT-4, ChatGPT 3.5, and various Llama and Gemma series models, finding that all state-of-the-art LLMs underperform in playing optimal poker. However, after fine-tuning, these models show marked improvements. We validate PokerBench by having models with different scores compete with each other, demonstrating that higher scores on PokerBench lead to higher win rates in actual poker games. Through gameplay between our fine-tuned model and GPT-4, we also identify limitations of simple supervised fine-tuning for learning optimal playing strategy, suggesting the need for more advanced methodologies for effectively training language models to excel in games. PokerBench thus presents a unique benchmark for a quick and reliable evaluation of the poker-playing ability of LLMs as well as a comprehensive benchmark to study the progress of LLMs in complex game-playing scenarios.

大模型扑克博弈策略评估微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。