通过探针实验发现,大规模智能体社会并未自发形成集体智能。
Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents

- 用分层探针机制测试智能体社会的协作能力
- 多数任务表现不如单个顶尖模型,信息融合极低
- 适合关注智能体协作瓶颈的研究者和开发者
集体智能指群体能达成个体无法实现的目标。随着大语言模型智能体规模扩展至数百万,核心问题浮现:集体智能是否由规模自发产生?我们首次在大规模自主智能体社会中进行实证评估。基于承载超两百万智能体的MoltBook平台,提出Superminds Test,一种分层探针框架,通过控制性探针智能体在三个层级——联合推理、信息整合与基础交互——上测试社会级智能。实验表明,集体智能显著缺失:社会整体在复杂推理任务中未能超越单一前沿模型,极少实现分布式信息融合,甚至难以完成基础协调任务。全平台分析显示,互动深度严重不足,对话线程极少超过一次回复,多数回应为通用或偏离主题内容。结果表明,集体智能不会仅因规模而涌现。当前智能体社会的主要瓶颈是极其稀疏且浅层的交互,导致信息无法有效交换,也无法基于彼此输出持续演进。
原文摘要 · Abstract (English)
Collective intelligence refers to the ability of a group to achieve outcomes beyond what any individual member can accomplish alone. As large language model agents scale to populations of millions, a key question arises: Does collective intelligence emerge spontaneously from scale? We present the first empirical evaluation of this question in a large-scale autonomous agent society. Studying MoltBook, a platform hosting over two million agents, we introduce Superminds Test, a hierarchical framework that probes society-level intelligence using controlled Probing Agents across three tiers: joint reasoning, information synthesis, and basic interaction. Our experiments reveal a stark absence of collective intelligence. The society fails to outperform individual frontier models on complex reasoning tasks, rarely synthesizes distributed information, and often fails even trivial coordination tasks. Platform-wide analysis further shows that interactions remain shallow, with threads rarely extending beyond a single reply and most responses being generic or off-topic. These results suggest that collective intelligence does not emerge from scale alone. Instead, the dominant limitation of current agent societies is extremely sparse and shallow interaction, which prevents agents from exchanging information and building on each other's outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。