研究自适应多智能体系统泛化能力,发现其表面表现好但实际协作机制失效。
Superficial Success vs. Internal Breakdown: An Empirical Study of Generalization in Adaptive Multi-Agent Systems

- 通过实证分析自适应多智能体系统的拓扑结构与协作行为
- 揭示其在跨领域任务中存在拓扑过拟合与虚假协同现象
- 适合关注多智能体系统可靠性与评估标准的研究者
自适应多智能体系统(MAS)被越来越多地用于解决复杂问题。然而,其优化过程的任务覆盖范围狭窄,引发对其能否作为通用系统使用的疑问。为填补这一空白,我们对自适应多智能体系统进行了广泛实证研究,揭示出两个关键发现:(1)拓扑过拟合——它们在不同领域间无法泛化;(2)虚假协同——虽然表面准确率合理,但底层智能体交互行为偏离理想多智能体系统模式,引发其实用性的担忧。这些发现凸显了在多智能体系统开发中优先考虑泛化能力的紧迫性,并推动评估协议超越简单的最终答案正确性。
原文摘要 · Abstract (English)
Adaptive multi-agent systems (MAS) are increasingly adopted to tackle complex problems. However, the narrow task coverage of their optimization raises the question of whether they can function as general-purpose systems. To address this gap, we conduct an extensive empirical study of adaptive MAS, revealing two key findings: (1) topological overfitting -- they fail to generalize across different domains; and (2) illusory coordination -- they achieve reasonable surface-level accuracy while the underlying agent interactions diverge from ideal MAS behavior, raising concerns about their practical utility. These findings highlight the pressing need to prioritize generalization in MAS development and motivate evaluation protocols that extend beyond simple final-answer correctness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。