构建大模型群体协作评估框架,揭示去中心化系统中的信息传播与协调瓶颈。
Benchmarking Emergent Coordination in Large-Scale LLM Populations: An Evaluation Framework on the MoltBook Archive
- 基于273万次交互数据,量化分析角色分化、信息扩散与协作任务表现
- 发现核心-边缘结构显著(轮廓系数0.91),信息传播呈重尾分布(α=2.57)
- 适用于研究大规模多智能体系统协同机制的科研人员与算法设计者
随着多智能体大语言模型系统规模扩大,评估其涌现的协作动态变得愈发关键。然而,现有评估范式多聚焦于单个智能体或小规模、显式结构化的群体,难以捕捉大规模去中心化群体中出现的自组织与病毒式信息传播现象。本文提出一种系统性评估框架,用于衡量开放环境中角色专业化、信息扩散及协作任务解决能力。我们在 MoltBook 观测档案库上验证该框架,该数据集包含 90,704 个自主智能体之间的 273 万次交互,建立了涌现协作的定量基准。评估结果揭示出明显的核心-边缘结构(轮廓系数 0.91)、重尾级联分布(α = 2.57),以及去中心化任务求解中严重的协调开销(相对于单智能体基线,Cohen's d = -0.88)。通过提供标准化评估任务与实证基准,本框架支持未来多智能体协议的严格比较,并使评估本身成为科学研究的对象。
原文摘要 · Abstract (English)
As multi-agent Large Language Model (LLM) systems scale, evaluating their emergent coordination dynamics becomes increasingly critical. However, current evaluation paradigms-focused on single agents or small, explicitly structured groups-fail to capture the self-organization and viral information dynamics that arise in large, decentralized populations. We introduce a systematic evaluation framework to benchmark role specialization, information diffusion, and cooperative task resolution in open agent environments. We demonstrate this framework on the MoltBook Observatory Archive, a dataset of 2.73M interactions among 90,704 autonomous agents, establishing quantitative baselines for emergent coordination. Our evaluation reveals a pronounced core-periphery structure (silhouette 0.91), heavy-tailed cascade distributions ($α= 2.57$), and severe coordination overhead in decentralized task resolution (Cohen's $d = -0.88$ against a single-agent baseline). By providing standardized evaluation tasks and empirical baselines, our framework enables the rigorous comparison of future multi-agent protocols and establishes evaluation itself as an object of scientific study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。