arXiv:2507.08616cs.MAcs.LG2025-07被引 30

提出新基准AgentsNet,评估多智能体系统在复杂网络中的协作与自组织能力。

AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs

  • 基于分布式系统与图论设计,测试多智能体协同策略形成能力。
  • 100个智能体的实验表明,前沿模型在小规模下表现好但规模扩大后性能下降。
  • 适用于研究大规模智能体协作的学者,尤其关注系统可扩展性者。

大型语言模型(LLMs)在多智能体系统中展现出强大的问题求解能力,但这类系统的自组织与协作效率仍存疑问。现有基准最多支持2-5个智能体,难以评估拓扑结构下的协同表现。为此,本文提出AgentsNet,一个基于分布式系统与图论经典问题的新基准,用于衡量多智能体系统在给定网络拓扑下的策略协同、自我组织与有效通信能力。我们评估了多种基线方法,包括需协商基础协议的同质智能体网络。结果发现,部分前沿大模型在小规模网络中表现良好,但当智能体数量增至100时性能显著下降。该基准可无限扩展,适配未来大模型的发展。

原文摘要 · Abstract (English)

Large-language models (LLMs) have demonstrated powerful problem-solving capabilities, in particular when organized in multi-agent systems. However, the advent of such systems also raises several questions on the ability of a complex network of agents to effectively self-organize and collaborate. While measuring performance on standard reasoning benchmarks indicates how well multi-agent systems can solve reasoning tasks, it is unclear whether these systems are able to leverage their topology effectively. Here, we propose AgentsNet, a new benchmark for multi-agent reasoning. By drawing inspiration from classical problems in distributed systems and graph theory, AgentsNet measures the ability of multi-agent systems to collaboratively form strategies for problem-solving, self-organization, and effective communication given a network topology. We evaluate a variety of baseline methods on AgentsNet including homogeneous networks of agents which first have to agree on basic protocols for organization and communication. We find that some frontier LLMs are already demonstrating strong performance for small networks but begin to fall off once the size of the network scales. While existing multi-agent benchmarks cover at most 2-5 agents, AgentsNet is practically unlimited in size and can scale with new generations of LLMs. As such, we also probe frontier models in a setup with up to 100 agents.

多智能体推理评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。