揭示复杂任务下多智能体系统优势,指明何时用多智能体更有效。
On the Importance of Task Complexity in Evaluating LLM-Based Multi-Agent Systems

- 从推理深度和能力广度双维度定义任务复杂度
- 发现多智能体比单智能体在高复杂度任务中优势更大
- 为设计智能体系统和评测基准提供理论依据
大型语言模型多智能体系统(LLM-MAS)为实现更高级的智能行为提供了有前景的范式。尽管近期研究显示某些任务中LLM-MAS优于单智能体系统(LLM-SAS),但缺乏系统的实验设计限制了结论的强度与普适性。本文认为,对任务复杂度(如所需序列推理长度和涉及能力多样性)的严谨理解,是评估LLM-MAS有效性的重要前提。为此,我们提出一个理论框架,将任务沿两个维度刻画:深度(代表推理长度)与宽度(代表能力多样性)。理论上分析了典型的多智能体辩论系统,并在具有不同深度与宽度的判别与生成任务上进行实证评估。理论与实证结果表明,LLM-MAS相较于LLM-SAS的优势随任务深度与宽度增加而提升,且深度的影响更为显著。该研究澄清了多智能体系统的适用场景,为未来方法设计与评测基准构建提供了原则性基础。
原文摘要 · Abstract (English)
Large language model multi-agent systems (LLM-MAS) offer a promising paradigm for harnessing collective intelligence to achieve more advanced forms of AI behaviour. While recent studies suggest that LLM-MAS can outperform LLM single-agent systems (LLM-SAS) on certain tasks, the lack of systematic experimental designs limits the strength and generality of these conclusions. We argue that a principled understanding of task complexity, such as the degree of sequential reasoning required and the breadth of capabilities involved, is essential for assessing the effectiveness of LLM-MAS in task solving. To this end, we propose a theoretical framework characterising tasks along two dimensions: depth, representing reasoning length, and width, representing capability diversity. We theoretically examine a representative class of LLM-MAS, namely the multi-agent debate system, and empirically evaluate its performance in both discriminative and generative tasks with varying depth and width. Theoretical and empirical results show that the benefit of LLM-MAS over LLM-SAS increases with both task depth and width, and the effect is more pronounced with respect to depth. This clarifies when LLM-MAS are beneficial and provides a principled foundation for designing future LLM-MAS methods and benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。