首个评估真实世界动态规划任务的多智能体基准,支持从简单到复杂的渐进式测试。
REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks
- 设计14个由简至繁的动态规划任务,含多智能体协作与突发干扰
- 支持三维度扩展:并行任务数、依赖复杂度、突发事件频率
- 覆盖主流大模型与框架,适合评估真实场景下的智能体系统
该基准套件为评估大语言模型和多智能体系统在真实世界规划与调度场景中的表现提供了全面框架。包含14个逐步升级的规划调度问题,涵盖多智能体协作、智能体间依赖关系及动态环境扰动等关键要素。每个问题可在三个维度上扩展:并行规划线程数、依赖关系复杂度、意外干扰频率,以实现对实时适应能力的考验。基准提供14个详细问题说明、15种对比方法(如随机策略、LPT、SPT、DRL-Liu等)、2项评估指标,以及基于GPT-4o、Claude-3.7、DeepSeek-R1等3+大模型和LangGraph、AutoGen、CrewAI、Swarm等4个现代框架的基线实现,支持单智能体与多智能体规划能力的严格测试。通过标准化评价标准与可扩展复杂度,该基准将向公众开放,推动更适应现实应用的可扩展、鲁棒性强的AI规划系统发展。
原文摘要 · Abstract (English)
This benchmark suite provides a comprehensive evaluation framework for assessing both individual LLMs and multi-agent systems in Real-world planning and scheduling scenarios. The suite encompasses 14 designed planning and scheduling problems that progress from basic to highly complex, incorporating key aspects such as multi-agent coordination, inter-agent dependencies, and dynamic environmental disruptions. Each problem can be scaled along three dimensions: the number of parallel planning threads, the complexity of inter-dependencies, and the frequency of unexpected disruptions requiring Real-time adaptation. The benchmark includes 14 detailed problem specifications, 15 comparison methods including Random, LPT, SPT, STPT, MPSR, DRL-Liu, GP, GEP, LSO, SPT/TWKR, DRL-Chen, DRL-Zhang, 2+ evaluation metrics, and baseline implementations using 3+ LLMs including GPT-4o, Claude-3.7, DeepSeek-R1, and 4 contemporary frameworks including LangGraph, AutoGen, CrewAI, and Swarm, enabling rigorous testing of both single-agent and multi-agent planning capabilities. Through standardized evaluation criteria and scalable complexity, this benchmark aims to be opened to public, and drive progress in developing more adaptable, robust, and scalable AI planning systems for Real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。