提出协作熵度量多大模型间的语义不确定性,提升协同决策可信度。
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
- 基于共享语义空间,融合模型内熵与模型间差异度,构建系统级不确定性度量。
- 在TriviaQA和SQuAD上优于传统方法,引入异构模型后效果提升更显著。
- 无需训练的后处理协调策略,适合部署于实际多模型协作系统。
多大模型系统中的不确定性估计仍以单模型为中心:现有方法仅量化各模型内部不确定性,未能充分捕捉模型间的语义分歧。为弥补这一空白,我们提出协作熵(CoE),一种统一的信息论度量,用于多大模型协作中的语义不确定性。CoE 定义于共享语义聚类空间,结合两个成分:模型内语义熵和模型间对集成均值的偏离程度。CoE 不是加权集成预测器,而是刻画协作信心与分歧的系统级不确定性度量。我们分析了CoE的若干核心性质,包括非负性、完全语义一致时为零、以及当单个模型坍缩为狄拉克分布时的行为。这些结果明确了降低模型内不确定性是否足够,以及残余模型间分歧何时依然存在。此外,我们提出一种简单的CoE引导、无需训练的后处理协调启发式方法作为实际应用。在使用LLaMA-3.1-8B-Instruct、Qwen-2.5-7B-Instruct和Mistral-7B-Instruct的TriviaQA和SQuAD上的实验表明,CoE在不确定性估计上优于标准熵与分歧基线,且随着引入更多异构模型,优势进一步扩大。总体而言,CoE为多大模型协作提供了有价值的不确定性感知视角。
原文摘要 · Abstract (English)
Uncertainty estimation in multi-LLM systems remains largely single-model-centric: existing methods quantify uncertainty within each model but do not adequately capture semantic disagreement across models. To address this gap, we propose Collaborative Entropy (CoE), a unified information-theoretic metric for semantic uncertainty in multi-LLM collaboration. CoE is defined on a shared semantic cluster space and combines two components: intra-model semantic entropy and inter-model divergence to the ensemble mean. CoE is not a weighted ensemble predictor; it is a system-level uncertainty measure that characterizes collaborative confidence and disagreement. We analyze several core properties of CoE, including non-negativity, zero-value certainty under perfect semantic consensus, and the behavior of CoE when individual models collapse to delta distributions. These results clarify when reducing per-model uncertainty is sufficient and when residual inter-model disagreement remains. We also present a simple CoE-guided, training-free post-hoc coordination heuristic as a practical application of the metric. Experiments on \textit{TriviaQA} and \textit{SQuAD} with LLaMA-3.1-8B-Instruct, Qwen-2.5-7B-Instruct, and Mistral-7B-Instruct show that CoE provides stronger uncertainty estimation than standard entropy- and divergence-based baselines, with gains becoming larger as additional heterogeneous models are introduced. Overall, CoE offers a useful uncertainty-aware perspective on multi-LLM collaboration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。