arXiv:2411.11057cs.AI2024-11

用复杂博弈游戏测试多智能体强化学习的策略与协作能力

Reinforcing Competitive Multi-Agents for Playing 'So Long Sucker'

  • 构建首个SLS游戏计算框架,支持智能体自对弈训练
  • 经典RL方法仅达理论最大奖励的一半,需约2000局训练
  • 适合研究联盟形成与策略欺骗的高级多智能体系统

本文将策略游戏So Long Sucker(SLS)作为多智能体强化学习(MARL)的新基准。与传统棋类或视频游戏不同,SLS具有联盟形成、策略欺骗和动态淘汰规则,对自主智能体构成独特挑战。我们首次发布SLS的公开计算框架,包含图形界面和算法评测支持。基于DQN、DDQN和Dueling DQN等经典深度强化学习方法,训练自对弈智能体学习游戏规则与基础策略。实验表明,尽管这些智能体达到理论最大奖励的大约一半,且持续优于随机基线,但仍需约2000局训练,且偶发非法操作,凸显经典强化学习在该场景下的潜力与局限。研究确立SLS为具备谈判意识的MARL基准,为未来融合博弈论推理、联盟感知策略及先进强化学习架构的研究开辟道路。

原文摘要 · Abstract (English)

This paper investigates the strategy game So Long Sucker (SLS) as a novel benchmark for multi-agent reinforcement learning (MARL). Unlike traditional board or video game testbeds, SLS is distinguished by its coalition formation, strategic deception, and dynamic elimination rules, making it a uniquely challenging environment for autonomous agents. We introduce the first publicly available computational framework for SLS, complete with a graphical user interface and benchmarking support for reinforcement learning algorithms. Using classical deep reinforcement learning methods (e.g., DQN, DDQN, and Dueling DQN), we train self-playing agents to learn the rules and basic strategies of SLS. Experimental results demonstrate that, although these agents achieve roughly half of the maximum attainable reward and consistently outperform random baselines, they require long training horizons (~2000 games) and still commit occasional illegal moves, highlighting both the promise and limitations of classical reinforcement learning. Our findings establish SLS as a negotiation-aware benchmark for MARL, opening avenues for future research that integrates game-theoretic reasoning, coalition-aware strategies, and advanced reinforcement learning architectures to better capture the social and adversarial dynamics of complex multi-agent games.

多智能体强化学习博弈策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。