arXiv:2411.05348cs.AI2024-11被引 14

为大模型量身定制星际2决策环境,支持全动作空间与多智能体协作。

LLM-PySC2: Starcraft II learning environment for Large Language Models

  • 构建首个支持大模型完整操作空间的星际2交互环境
  • 大模型在复杂场景中可获胜但决策不稳定,尤其在多智能体设置下
  • 适合研究大模型决策、多智能体协作与强化学习融合的学者

大语言模型(LLMs)在智能决策任务中展现出巨大潜力,覆盖从游戏AI到复杂战略规划的多种应用。然而,过去十年广泛用于验证决策算法的星际2平台,尚未为这一新兴领域提供有效支持。由于LLMs无法直接对接pysc2后端数百个动作,且缺乏原生多智能体(MA)协作支持,我们提出LLM-PySC2环境。该环境首次为大模型提供完整的pysc2动作空间、多模态信息及游戏维基知识。通过异步查询架构,环境能保持恒定延迟,不受智能体数量影响。实验评估了大模型在宏观决策和微观操作场景中的表现,涵盖传统SMAC任务及新提出的任务。结果表明,大模型具备在复杂场景中取胜的潜力,但无法持续生成正确决策,尤其在还原后的pysc2动作空间和多智能体设置下表现不佳。缺乏任务相关指令时,预训练模型易出现幻觉和协作效率低下问题。研究揭示,星际2在大模型时代仍具挑战性,亟需发展更先进的大模型决策系统,而本研究所提出的LLM-PySC2环境将推动未来基于大模型的决策解决方案发展。

原文摘要 · Abstract (English)

The tremendous potential has been demonstrated by large language models (LLMs) in intelligent decision-making problems, with unprecedented capabilities shown across diverse applications ranging from gaming AI systems to complex strategic planning frameworks. However, the StarCraft II platform, which has been widely adopted for validating decision-making algorithms in the past decade, has not yet provided substantial support for this emerging domain. To address issues that LLMs cannot interface with the hundreds of actions of the pysc2 backend and the lack of native support for multi-agent (MA) collaboration, we propose the LLM-PySC2 environment. This is the first environment that offers LLMs the complete pysc2 action space with sufficient multi-modal information and game Wiki knowledge. With an asynchronous query architecture, the environment efficiently interacts with LLMs that maintain a constant latency regardless of the scale of the agents' population. In the experiments, we evaluated LLMs' decision-making performance in both the macro-decision and micro-operation scenarios, with traditional StarCraft II Multi-Agent Challenge (SMAC) tasks and a series of new proposed. Results indicate that LLMs possess the potential to achieve victories in complex scenarios but cannot constantly generate correct decisions, especially in the recovered pysc2 action space and MA settings. Without task-relevant instructions, the pre-trained models suffer from issues such as hallucinations and inefficient collaboration. Our findings suggest that StarCraft II still challenges in the era of large models, revealing that there is a lot to do to develop an advanced LLM decision-making system, and the proposed LLM-PySC2 environment will support future development of LLM-based decision-making solutions.

大模型决策星际2多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。