开源麻将类游戏基准测试,挑战多智能体决策新高度。
OpenGuanDan: A Large-Scale Imperfect Information Game Benchmark
- 构建四人回合制中国牌类游戏仿真环境,支持复杂策略博弈
- 学习型智能体胜过规则型但未达超人类水平,暴露现有方法短板
- 支持人机协作与大模型接入,适合多智能体强化学习研究者
数据驱动的人工智能发展依赖大规模基准测试。尽管近年来在棋类、牌类和电子竞技等领域取得显著进展,仍需更具挑战性的基准以推动研究。本文提出 OpenGuanDan,一个新型基准,可高效模拟四人多轮中国牌类游戏「掼蛋」,并全面评估基于学习和基于规则的掼蛋智能体。该基准包含不完全信息、庞大信息集与动作空间、合作与竞争混合目标、长时序决策、可变动作空间及动态组队等复杂特性,构成对现有智能决策方法的严峻考验。每个玩家独立接口支持人机交互与大语言模型集成。实验上进行两类评估:(1)所有掼蛋智能体间的两两对抗;(2)人机对战。结果表明,当前学习型智能体虽显著优于规则型,但仍未能达到超人类表现,凸显多智能体智能决策领域亟待深入研究。项目已公开于 https://github.com/GameAI-NJUPT/OpenGuanDan。
原文摘要 · Abstract (English)
The advancement of data-driven artificial intelligence (AI), particularly machine learning, heavily depends on large-scale benchmarks. Despite remarkable progress across domains ranging from pattern recognition to intelligent decision-making in recent decades, exemplified by breakthroughs in board games, card games, and electronic sports games, there remains a pressing need for more challenging benchmarks to drive further research. To this end, this paper proposes OpenGuanDan, a novel benchmark that enables both efficient simulation of GuanDan (a popular four-player, multi-round Chinese card game) and comprehensive evaluation of both learning-based and rule-based GuanDan AI agents. OpenGuanDan poses a suite of nontrivial challenges, including imperfect information, large-scale information set and action spaces, a mixed learning objective involving cooperation and competition, long-horizon decision-making, variable action spaces, and dynamic team composition. These characteristics make it a demanding testbed for existing intelligent decision-making methods. Moreover, the independent API for each player allows human-AI interactions and supports integration with large language models. Empirically, we conduct two types of evaluations: (1) pairwise competitions among all GuanDan AI agents, and (2) human-AI matchups. Experimental results demonstrate that while current learning-based agents substantially outperform rule-based counterparts, they still fall short of achieving superhuman performance, underscoring the need for continued research in multi-agent intelligent decision-making domain. The project is publicly available at https://github.com/GameAI-NJUPT/OpenGuanDan.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。