AI在诈唬扑克中击败人类高手,靠自对弈强化学习。
Outbidding and Outbluffing Elite Humans: Mastering Liar's Poker via Self-Play and Reinforcement Learning
- 用自对弈和深度强化学习训练AI,不依赖预设策略。
- 在多人对局中胜率超50%,赢钱金额领先人类和大模型。
- 策略随机性强,难被人类高手利用,适合博弈研究者。
AI研究者长期将扑克类游戏作为多玩家动态、信息不完全和不确定性推理的测试平台。尽管近期突破已使AI在无限注德州扑克中达到顶尖人类水平,但多数对局仅限两人参与,多玩家互动有限。本文提出Solly,首个在简化版诈唬扑克中达到人类顶级水平的AI代理。采用无模型、基于演员-评论家的深度强化学习算法进行自对弈训练。在一对一及多人对局中,Solly的胜率超过50%,所获权益(赢钱金额)均表现优异,且超越大型语言模型(包括具备推理能力的模型)。该模型发展出新颖的出价策略,有效实现随机化行动,且不易被世界级人类玩家 exploit。实验验证其在多玩家环境中具备卓越的持续参与能力。
原文摘要 · Abstract (English)
AI researchers have long focused on poker-like games as a testbed for environments characterized by multi-player dynamics, imperfect information, and reasoning under uncertainty. While recent breakthroughs have matched elite human play at no-limit Texas hold'em, the multi-player dynamics are subdued: most hands converge quickly with only two players engaged through multiple rounds of bidding. In this paper, we present Solly, the first AI agent to achieve elite human play in reduced-format Liar's Poker, a game characterized by extensive multi-player engagement. We trained Solly using self-play with a model-free, actor-critic, deep reinforcement learning algorithm. Solly played at an elite human level as measured by win rate (won over 50% of hands) and equity (money won) in heads-up and multi-player Liar's Poker. Solly also outperformed large language models (LLMs), including those with reasoning abilities, on the same metrics. Solly developed novel bidding strategies, randomized play effectively, and was not easily exploitable by world-class human players.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。