多智能体大模型在分散信息下集体推理能力差,关键问题在于无法察觉他人未知信息。
Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs
- 构建隐含信息任务基准HiddenBench,分离集体与个体推理能力。
- 多智能体平均准确率仅30.1%,远低于单智能体80.7%的完整信息表现。
- 引入轻量结构化通信可显著提升跨模型群体推理效果,适合系统评估研究者。
基于大语言模型的多智能体系统本应通过整合分布式信息提升决策能力,但其实际表现尚无系统性评估。我们提出HiddenBench,一个基于隐含议题范式的65项任务基准,将集体推理中的分布式信息处理能力与个体推理能力分离。在15个前沿大模型上测试发现,多智能体在分布式信息下的平均准确率为30.1%,而单智能体在完整信息下可达80.7%。该差距源于系统性失效:智能体无法识别或应对潜在的信息不对称,即未能推理他人所知但未表达的内容,导致过早收敛于共享证据,而关键分布式事实被忽略。此问题在不同提示策略、通信深度和群体规模下均存在,且随群体扩大加剧。尽管部分模型(如Gemini-2.5-Flash/Pro)表现更优,但模型规模与个体推理准确率无法预测集体表现。进一步表明,这一瓶颈可通过轻量级结构化通信协议有效缓解。研究揭示了多智能体大模型在集体信息探索中的核心缺陷,并提供了理论支撑、可复现的诊断框架。
原文摘要 · Abstract (English)
Multi-agent systems built on large language models (LLMs) are expected to enhance decision-making by pooling distributed information, yet systematically evaluating this capability has remained challenging. We introduce HiddenBench, a 65-task benchmark grounded in the Hidden Profile paradigm, which isolates collective reasoning under distributed information from individual reasoning ability. Evaluating 15 frontier LLMs, we find that multi-agent LLMs achieve only 30.1% accuracy under distributed information, compared to 80.7% accuracy for single agents given complete information. We trace this gap to a systematic failure mode: agents cannot recognize or act under latent information asymmetry -- they fail to reason about what others might know but have not yet expressed, leading to premature convergence on shared evidence while critical distributed facts remain unexplored. These failures persist across prompting strategies, communication depths, and group sizes -- and worsen as groups scale. While some models (e.g., Gemini-2.5-Flash/Pro) outperform others, neither model scale nor individual reasoning accuracy reliably predicts collective performance. We further show that this bottleneck is actionable: a lightweight structured communication protocol substantially improves collective reasoning across model families. Our results identify failures in collective information exploration in decision-making as a key limitation of multi-agent LLMs, and provide a theory-grounded, reproducible framework for diagnosing collective reasoning failures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。