提出反事实图框架,让多智能体LLM更准确评估共识可信度。
Counterfactual Graph for Multi-Agent LLM Calibration

- 通过对比有通信与无通信的智能体图结构,捕捉失败相关性。
- 在五个基准上提升可靠性判别能力,误差率降低12%-23%。
- 适合需要高可靠决策的多智能体系统,如自动驾驶协同决策。
多智能体大模型常将共识视为可靠证据:当多个智能体给出相同答案时,认为该答案更可信。我们发现,智能体间通信可能导致错误共识和关联性失败,使得相同的投票比例在不同拓扑中可能反映真实一致或过度自信。为此,我们提出CAGE-CAL——一种基于反事实图的多智能体大模型校准框架。对每个问题,CAGE-CAL比较实际通信后的智能体图与对应的无通信反事实图,同时捕获成对失败相关性和群体依赖关系。不单纯统计共识数量,而是估计观察到与无通信状态间的依赖性变化,据此校准置信度。在五个基准测试中,CAGE-CAL显著提升可靠性判别能力,且校准后置信度进一步优于最优固定拓扑策略。
原文摘要 · Abstract (English)
Multi-agent LLM systems often treat agreement as evidence: when many agents in a panel give the same answer, that answer is assumed to be more reliable. We show that this assumption can fail after agents communicate. Communication can induce correlated failures and false consensus, so the same vote share may reflect reliable agreement in one topology but over-confidence in another. We propose CAGE-CAL, a counterfactual agent-graph calibration framework for multi-agent LLMs. For each query, CAGE-CAL compares an observed post-communication agent graph with a matched counterfactual no-communication graph, capturing both pairwise failure correlations and group-level dependencies. Rather than simply counting how many agents agree, CAGE-CAL estimates the counterfactual shift between observed and no-communication dependence, and calibrates confidence accordingly. Across five benchmarks, CAGE-CAL improves reliability discrimination with competitive ECE, and its calibrated confidence further improves topology selection over the best fixed-topology strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。