揭示大模型多智能体规划的可靠性上限,指出通信限制导致性能不如中心化决策。
On the Reliability Limits of LLM-Based Multi-Agent Planning
- 将多智能体规划建模为有限无环决策网络,分析信息传递与压缩机制。
- 在通信预算下,性能损失可量化为后验分布差异或条件互信息。
- 实验验证理论,适合研究大模型协作与可信决策的学者参考。
本文研究基于大语言模型的多智能体规划作为委托决策问题的可靠性极限。我们将该架构建模为有限无环决策网络,其中多个阶段处理共享模型上下文信息,通过容量受限的语言接口通信,并可能触发人工审查。结果表明,在无外部新信号的情况下,任何委托网络在决策理论上均劣于拥有相同信息的集中式贝叶斯决策者。在公共证据场景下,优化有限通信预算下的多智能体有向无环图,等价于在共享信号上选择一个预算受限的随机实验。我们还刻画了通信与信息压缩带来的损失:在合理评分规则下,集中式贝叶斯价值与通信后价值之差可表示为期望后验发散,对数损失下退化为条件互信息,布里尔分数下退化为期望平方后验误差。这些结果刻画了委托式大模型规划的根本可靠性极限。在受控问题集上的大模型实验进一步验证了这些结论。
原文摘要 · Abstract (English)
This technical note studies the reliability limits of LLM-based multi-agent planning as a delegated decision problem. We model the LLM-based multi-agent architecture as a finite acyclic decision network in which multiple stages process shared model-context information, communicate through language interfaces with limited capacity, and may invoke human review. We show that, without new exogenous signals, any delegated network is decision-theoretically dominated by a centralized Bayes decision maker with access to the same information. In the common-evidence regime, this implies that optimizing over multi-agent directed acyclic graphs under a finite communication budget can be recast as choosing a budget-constrained stochastic experiment on the shared signal. We also characterize the loss induced by communication and information compression. Under proper scoring rules, the gap between the centralized Bayes value and the value after communication admits an expected posterior divergence representation, which reduces to conditional mutual information under logarithmic loss and to expected squared posterior error under the Brier score. These results characterize the fundamental reliability limits of delegated LLM planning. Experiments with LLMs on a controlled problem set further demonstrate these characterizations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。