新基准HLSMAC挑战高阶战略决策,12个经典兵法场景评测智能体全局思维。
HLSMAC: A New StarCraft Multi-Agent Challenge for High-Level Strategic Decision-Making
- 基于三十六计设计12个星战2合作场景,聚焦高层战略而非操作细节。
- 引入能力利用率、推进效率等新指标,超越单纯胜率评估综合表现。
- 适合研究战略规划、多智能体协作与大模型决策的学者使用。
基准测试对评估多智能体强化学习(MARL)算法至关重要。尽管基于《星际争霸Ⅱ》的环境推动了MARL显著进展,但现有基准如SMAC主要关注微操,限制了对高层战略智能的全面评估。为此,我们提出HLSMAC,一个包含12个精心设计的《星际争霸Ⅱ》合作场景的新基准,其灵感来自《三十六计》中的经典战术策略。每个场景对应一种特定计谋,旨在挑战智能体在战术机动、时机协调和欺骗等方面的多样化战略能力,从而为高阶战略决策能力的评估开辟新路径。我们还提出了超越传统胜率的多维度新指标,如技能使用率和推进效率,以评估智能体在HLSMAC环境中的整体表现。我们将前沿MARL算法与基于大语言模型(LLM)的智能体集成到该基准中并进行系统实验。结果表明,HLSMAC是一个强大的测试平台,能有效推动多智能体战略决策的发展。
原文摘要 · Abstract (English)
Benchmarks are crucial for assessing multi-agent reinforcement learning (MARL) algorithms. While StarCraft II-related environments have driven significant advances in MARL, existing benchmarks like SMAC focus primarily on micromanagement, limiting comprehensive evaluation of high-level strategic intelligence. To address this, we introduce HLSMAC, a new cooperative MARL benchmark with 12 carefully designed StarCraft II scenarios based on classical stratagems from the Thirty-Six Stratagems. Each scenario corresponds to a specific stratagem and is designed to challenge agents with diverse strategic elements, including tactical maneuvering, timing coordination, and deception, thereby opening up avenues for evaluating high-level strategic decision-making capabilities. We also propose novel metrics across multiple dimensions beyond conventional win rate, such as ability utilization and advancement efficiency, to assess agents' overall performance within the HLSMAC environment. We integrate state-of-the-art MARL algorithms and LLM-based agents with our benchmark and conduct comprehensive experiments. The results demonstrate that HLSMAC serves as a robust testbed for advancing multi-agent strategic decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。