为群体AI建立科学框架,用新指标区分协作真进步与资源堆叠。
Towards a Science of Collective AI: LLM-based Multi-Agent Systems Need a Transition from Blind Trial-and-Error to Rigorous Science
- 提出协作增益指标Γ,剥离资源增加对效果的干扰。
- 构建因素库,系统识别影响协作的关键设计因子。
- 适合想从试错转向可复现研究的AI系统设计者。
大语言模型(LLM)的进展极大拓展了多智能体系统(MAS)的能力,在复杂开放领域展现出显著成效。然而,该领域仍严重依赖经验性的试错方法,缺乏统一、严谨的科学框架来实现系统性优化。这一瓶颈源于归因模糊:一是缺乏结构化的因素分类,导致调整无方向;二是缺少统一评估指标,无法区分真正的协作收益与单纯资源积累。本文倡导通过集成框架实现设计科学的转型,提出以协作增益指标Γ作为科学标准,分离内在协作优势与预算增加带来的影响。基于Γ,我们构建因子归因范式,系统识别驱动协作的核心因素。为此,我们建立了系统的MAS因子库,将设计空间划分为控制层级预设和信息层级动态。该框架推动群体AI从盲目实验迈向严谨科学,为真正意义上的集体智能科学铺路。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have greatly extended the capabilities of Multi-Agent Systems (MAS), demonstrating significant effectiveness across a wide range of complex and open-ended domains. However, despite this rapid progress, the field still relies heavily on empirical trial-and-error. It lacks a unified and principled scientific framework necessary for systematic optimization and improvement. This bottleneck stems from the ambiguity of attribution: first, the absence of a structured taxonomy of factors leaves researchers restricted to unguided adjustments; second, the lack of a unified metric fails to distinguish genuine collaboration gain from mere resource accumulation. In this paper, we advocate for a transition to design science through an integrated framework. We advocate to establish the collaboration gain metric ($Γ$) as the scientific standard to isolate intrinsic gains from increased budgets. Leveraging $Γ$, we propose a factor attribution paradigm to systematically identify collaboration-driving factors. To support this, we construct a systematic MAS factor library, structuring the design space into control-level presets and information-level dynamics. Ultimately, this framework facilitates the transition from blind experimentation to rigorous science, paving the way towards a true science of Collective AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。