构建大规模真实火灾救援模拟,测试多智能体协作能力
CREW-WILDFIRE: Benchmarking Agentic Multi-Agent Collaborations at Scale
- 基于人类-智能体协同平台,生成动态火灾场景
- 在大地图、部分观测下暴露长时规划与通信短板
- 适合研究可扩展多智能体系统的研究人员
尽管大型语言模型(LLM)驱动的多智能体系统发展迅速,现有基准在评估其可扩展性、鲁棒性和协调能力方面仍显不足。当前环境多聚焦于小规模、全可观测或低复杂度领域,难以支撑下一代多智能体自主智能框架的研发与评估。我们提出CREW-Wildfire,一个开源基准,基于人类-智能体协同平台CREW构建,提供程序化生成的野火响应场景,包含大地图、异构智能体、部分可观测性、随机动态及长周期规划目标。该环境支持通过模块化感知与执行模块实现低层控制与高层自然语言交互。我们实现了多个前沿的基于LLM的多智能体自主智能框架,并揭示显著性能差距,凸显在不确定性下大规模协调、通信、空间推理与长期规划中的未解挑战。通过提供更真实的复杂性、可扩展架构和行为评估指标,CREW-Wildfire为推动可扩展多智能体自主智能研究奠定关键基础。所有代码、环境、数据与基线将公开,以支持该新兴领域的后续研究。
原文摘要 · Abstract (English)
Despite rapid progress in large language model (LLM)-based multi-agent systems, current benchmarks fall short in evaluating their scalability, robustness, and coordination capabilities in complex, dynamic, real-world tasks. Existing environments typically focus on small-scale, fully observable, or low-complexity domains, limiting their utility for developing and assessing next-generation multi-agent Agentic AI frameworks. We introduce CREW-Wildfire, an open-source benchmark designed to close this gap. Built atop the human-AI teaming CREW simulation platform, CREW-Wildfire offers procedurally generated wildfire response scenarios featuring large maps, heterogeneous agents, partial observability, stochastic dynamics, and long-horizon planning objectives. The environment supports both low-level control and high-level natural language interactions through modular Perception and Execution modules. We implement and evaluate several state-of-the-art LLM-based multi-agent Agentic AI frameworks, uncovering significant performance gaps that highlight the unsolved challenges in large-scale coordination, communication, spatial reasoning, and long-horizon planning under uncertainty. By providing more realistic complexity, scalable architecture, and behavioral evaluation metrics, CREW-Wildfire establishes a critical foundation for advancing research in scalable multi-agent Agentic intelligence. All code, environments, data, and baselines will be released to support future research in this emerging domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。