提出新方法量化大模型多智能体系统中的错误传播风险。
PropUQ-MAS: Propagation-Aware Uncertainty Quantification for LLM Multi-Agent Systems

- 将多智能体交互建模为通信图,追踪错误传播路径。
- 实验显示AUROC提升6.10%,PRR提升47.58%。
- 适合关注多智能体系统可靠性与可信推理的研究者。
基于大语言模型的多智能体系统(MAS)通过角色专业化智能体间的协作解决复杂任务。然而,智能体间的依赖关系引入了超越单个智能体故障的可靠性风险,例如中间消息中的错误可能被下游智能体继承并放大。现有不确定性量化(UQ)方法主要针对独立响应或单智能体推理,难以捕捉多智能体系统中的不确定性传播。为此,我们提出PropUQ-MAS,一种考虑错误传播的UQ框架,将多智能体执行过程建模为通信结构图,通过结合局部不确定性与上游消息继承的不确定性来估计每一步的可靠性。大量实验表明,PropUQ-MAS在多智能体系统中持续提升了不确定性量化效果,平均相对增益达AUROC +6.10%、PRR +47.58%。
原文摘要 · Abstract (English)
LLM-based multi-agent systems (MAS) solve complex tasks through communication among role-specialized agents. However, inter-agent dependencies introduce reliability risks beyond isolated agent failures. For instance, errors in intermediate messages could be inherited and amplified by downstream agents. Existing uncertainty quantification (UQ) methods mainly target isolated responses or single-agent reasoning, and therefore fail to capture uncertainty propagation in MAS. To this end, we propose PropUQ-MAS, an error propagation-aware UQ framework that represents MAS execution as a communication-structured graph and estimates each step's reliability by combining local uncertainty with uncertainty inherited from upstream messages. Extensive experiments demonstrate that PropUQ-MAS consistently improves UQ in MAS, with average relative gains of +6.10% in AUROC and +47.58% in PRR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。