arXiv:2606.13591cs.AIcs.LG2026-06

为多智能体系统设计统一置信度评估方法,提升决策可靠性。

Multiagent Protocols with Aggregated Confidence Signals

论文配图:Multiagent Protocols with Aggregated Confidence Signals
图 1 · 摘自论文原文
  • 将不同模型的原始置信度转换为可比信号,通过软投票或贝叶斯融合生成系统级置信度。
  • 系统置信度在AUARC指标上显著优于单个模型或传统辩论基线,任务准确率保持稳定。
  • 适用于多种模型组合与任务类型,尤其适合需要可靠置信评估的复杂决策场景。

置信度在自然语言处理中用于可靠性判断与下游决策,但现有方法无法为多智能体系统生成或评估置信度。以往研究仅在智能体辩论中使用置信度加权消息、触发辩论或校准个体,却未将其聚合为系统整体置信度。本文提出三种协议:先将原始置信度信号转换为可比形式,再通过软投票或称为贝叶斯融合的概率融合方式合并,生成单一系统置信度。该置信度在AUARC指标上显著优于最优单个模型或标准辩论基线,同时在准确率(F1得分)上保持稳定,并恢复了传统辩论在更模糊任务中的性能损失。通过分析序列概率与自报告两种估计器,结合参数与非参数校准器,发现校准对两类估计器均提升F1,而AUARC对校准依赖较小。在五个基准、四种任务类型下,评估六组同质与异质辩论对,覆盖多种模型能力与规模。

原文摘要 · Abstract (English)

Confidence is used for reliability, oversight, and a range of downstream decision tasks in Natural Language Processing (NLP), yet no existing method produces or evaluates a confidence for the output of a multiagent system. Prior work uses confidence within multiagent debate (MAD) to weight messages, trigger debate, or calibrate individual agents, but it never aggregates these into a single confidence for the system itself. We introduce three protocols that produce a final answer along with a single aggregated confidence by first transforming raw confidence signals to make them comparable across models, then combining them via soft voting or a probability fusion we call Bayesian fusion. This aggregated confidence is substantially more discriminative (AUARC) than that of the best single agent or the standard debate baselines, while correctness (F1-score) stays stable and recovers the losses MAD incurs on more ambiguous tasks. Analyzing two estimators, sequence probability and self-report, alongside parametric and non-parametric calibrators, we find that calibration improves F1 for both estimators while AUARC is less reliant on it. We evaluate six homogeneous and heterogeneous debating pairs per benchmark, across five benchmarks and four task types, spanning a range of model capabilities and sizes.

多智能体置信度评估决策融合校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。