arXiv:2509.06235cs.AIcs.MA2025-09被引 5

评测大模型智能体在竞争性团队环境中的表现,推动多智能体博弈研究。

PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments

  • 构建Minecraft实时对抗赛场景,支持多轮测试与规则对手
  • 自研TactiCrafter智能体实现战术协作与对手策略适应
  • 开源评估框架,适合研究多智能体博弈与自适应学习

基于大模型的智能体在协作与战略推理任务中展现出潜力,但在竞争性多智能体环境中的表现仍缺乏研究。为此,我们提出PillagerBench,一个在Minecraft中实现实时对抗团队赛的新型评估框架,提供可扩展API、多轮测试及基于规则的内置对手,确保公平、可复现的比较。我们还提出了TactiCrafter,一种基于大模型的多智能体系统,通过人类可读战术促进团队协作,学习因果依赖关系,并适应对手策略。评估显示,TactiCrafter优于基线方法,且通过自我对弈展现自适应学习能力。我们还分析了其在多个游戏回合中的学习过程与策略演化。为促进进一步研究,我们已开源PillagerBench,推动竞争性多智能体人工智能的发展。

原文摘要 · Abstract (English)

LLM-based agents have shown promise in various cooperative and strategic reasoning tasks, but their effectiveness in competitive multi-agent environments remains underexplored. To address this gap, we introduce PillagerBench, a novel framework for evaluating multi-agent systems in real-time competitive team-vs-team scenarios in Minecraft. It provides an extensible API, multi-round testing, and rule-based built-in opponents for fair, reproducible comparisons. We also propose TactiCrafter, an LLM-based multi-agent system that facilitates teamwork through human-readable tactics, learns causal dependencies, and adapts to opponent strategies. Our evaluation demonstrates that TactiCrafter outperforms baseline approaches and showcases adaptive learning through self-play. Additionally, we analyze its learning process and strategic evolution over multiple game episodes. To encourage further research, we have open-sourced PillagerBench, fostering advancements in multi-agent AI for competitive environments.

多智能体大模型游戏AI自适应学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。