arXiv:2606.02646physics.soc-phcs.AI2026-06被引 1

发现多智能体大模型团队规模存在有效上限,且可用简单公式预测。

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size

  • 提出一个描述团队有效规模的双参数缩放定律,揭示团队效能随人数增长的三种模式。
  • 30个密集讨论的智能体在难题上效果等同于1个,答案多样性几乎不增。
  • 仅异构团队能突破性能天花板,通信方式或重复思考无法提升效率。

推理时多智能体大模型缺乏统一度量单位:名义上的智能体数量混淆了成本与独立证据。我们推导出一个双参数缩放律 $R(N) = N_ ext{eff}/N = 1/(1+c(N-1)N^{-β})$,其中指数 $β$ 将任一配置划分为三类渐近行为——硬性上限($β=0$)、次线性增长($0<β<1$)或线性增长($β≥1$)。均值场定理表明,同伴数 $k$ 与辩论轮数 $τ$ 仅通过乘积 $kτ$ 影响动态过程。该规律在答案多样性与正确性冗余两个层面均适用。在44个(模型×任务×条件)组合中,涵盖同伴辩论、自修正、随机噪声对照、自一致性、三个开源模型家族(Qwen, Llama, Ministral)从7B到32B规模,以及前沿API(Gemini),思考型模型、异构团队和稀疏通信,所有条件下函数形式拟合均达 $R^2 > 0.99$,仅参数 $(c, β)$ 变化。在自由形式数学任务中,密集同伴影响使答案层面的缩放从次线性退化为硬性上限;正确性层面始终维持硬性上限。三项发现具实际意义:(i) 30个密集辩论智能体在MMLU-Hard上答案多样性不及单个智能体;(ii) 噪声对照组在自由形式数学任务上与自修正表现一致,且在4倍规模下仍匹配,说明‘辩论’带来的增益主要源于重审而非同伴内容;(iii) 单个 $N ≤ 5$ 的先导实验可预测 $N=30$ 的结构上限;在测试范围内,唯有架构异质性(异构团队)能降低 $c$ 并逃离硬性上限,通信模式干预无效。

原文摘要 · Abstract (English)

Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence. We derive a two-parameter scaling law $R(N) = N_\text{eff}/N = 1/(1+c(N-1)N^{-β})$ where the regime exponent $β$ classifies any configuration into one of three asymptotic regimes -- hard-ceiling at $1/c$ ($β= 0$), sublinear at $N^β/c$ ($0 < β< 1$), or linear ($β\ge 1$), and a mean-field theorem predicts that peer count $k$ and rounds $τ$ during agent debate enter the dynamics only through their product $kτ$. The law applies at two levels: answer diversity and correctness redundancy. Across 44 (model $\times$ task $\times$ condition) cells spanning peer debate, self-correction, random-noise placebo, self-consistency, three open-weight families (Qwen, Llama, Ministral) at scales from 7B to 32B with a frontier API check (Gemini), thinking models, heterogeneous teams, and sparse communication, the functional form fits every condition at $R^2 > 0.99$; only $(c, β)$ shifts. On free-form math, dense peer influence collapses the answer-level regime from sublinear into hard-ceiling; correctness-level fits remain hard-ceiling throughout. Three findings have practical implications. \emph{(i)}~Thirty dense debating agents produce no more answer diversity than one on MMLU-Hard. \emph{(ii)}~A noise placebo tracks self-correction on free-form math and at $4\times$ scale, so within homogeneous teams the gain commonly attributed to ``debate'' comes from re-evaluation, not peer content. \emph{(iii)}~A single $N \le 5$ pilot predicts the $N=30$ structural ceiling, and within the configurations tested only architectural diversity (heterogeneous teams) lowers $c$ and escapes the hard-ceiling regime, communication-mode interventions do not.

多智能体缩放定律团队效率大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。