arXiv:2608.20564cs.AI2026-08

让多个大模型协作时自动判断何时该说、谁该说,还能保证决策靠谱。

Consilience: Conformally Calibrated Communication Control for Hidden-Profile Multi-Agent Reasoning

论文配图:Consilience: Conformally Calibrated Communication Control for Hidden-Profile Multi-Agent Reasoning
图 1 · 摘自论文原文
  • 基于不确定性和分歧度动态决定每轮对话该谁发言、说什么。
  • 在12个任务上准确率超固定和自由讨论模式,甚至优于全知信息基线。
  • 首次实现对话行为的数学可信保障,适合高可靠场景的多智能体系统。

多智能体大模型系统可通过整合不同视角提升推理能力,但其效果依赖于通信协调,尤其在隐藏信息场景中——每个代理仅掌握部分证据。现有协议(如固定轮次、轮流发言、无结构辩论)无法保证对话行为的合理性。本文提出 Consilience,一个推理时的协同控制框架,可在分布式私有信息下引导并验证多智能体通信。每轮中,Consilience 用紧凑状态表征不确定性、分歧、证据增益、冗余与过早共识,进而选择沟通干预(质疑、澄清、索证或路由)及合适发言人。核心贡献是逐轮的共形校准机制,提供无需分布假设、有限样本下的保证:在达到某轮的前提下,控制器单步后悔值以至少 1−α 的边际概率被校准阈值约束;接受机制通过替换不合格提议,确保执行动作同样满足该保证。在涵盖12个开放与闭源权重语言模型的 HiddenBench 隐藏信息任务上,Consilience 在决策准确率和通信效率上均优于固定与无结构协议,有时甚至超越所有代理皆全知的基线。结果表明,经过认证的自适应通信控制比增加信息可见性更具价值,为多智能体大模型协作提供了可信赖的实用机制。

原文摘要 · Abstract (English)

Multi-agent LLM systems can improve reasoning by pooling diverse perspectives, but their effectiveness depends on coordinating communication, particularly in hidden-profile settings where each agent holds only part of the evidence required for a correct decision. Existing protocols, including fixed schedules, round-robin exchange, and unstructured debate, provide no guarantee that a conversational action is appropriate. We propose Consilience, an inference-time orchestration framework that both steers and certifies multi-agent communication under distributed private information. At each turn, Consilience summarizes the discussion using a compact state capturing uncertainty, disagreement, evidence gain, redundancy, and premature consensus, then selects both a communication intervention (challenge, clarify, seek evidence, or route) and an appropriate speaker. Its central contribution is a round-wise conformal calibration procedure that provides a distribution-free, finite-sample guarantee: at each discussion round, conditional on reaching that round, the one-step regret of a controller's proposed action is bounded by a calibrated threshold with marginal probability at least 1 - alpha; an acceptance mechanism enforces the same guarantee for the executed action by replacing inadmissible proposals. On HiddenBench-style hidden-profile tasks spanning 12 open and closed weight language models, Consilience improves decision accuracy and communication efficiency over fixed and unstructured discussion protocols, sometimes surpassing a full-information baseline where every agent observes all evidence. These results demonstrate that certified adaptive communication control can be more valuable than increasing information availability, providing a practical mechanism for reliable multi-agent LLM coordination.

多智能体推理增强可信通信共形校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。