arXiv:2607.20529cs.LGcs.AI2026-07

用可信度评估提升多大模型系统的预测可靠性

Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement

论文配图:Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement
图 1 · 摘自论文原文
  • 基于决策理论设计可信度评估机制,通过校准问题衡量模型不确定性
  • 在异构和污染环境下,该方法准确率显著优于传统平均法
  • 适合需要高可靠性的多模型集成场景,如安全关键应用

大型语言模型(LLM)集成正被广泛用于提升预测可靠性,但现有融合方法通常假设所有模型可信度相同,忽视了其不确定性质量的差异。这一假设在异构模型中表现不佳,导致集成结果易受不可靠或恶意模型影响。本文将多模型聚合建模为不确定性感知的可信度估计问题,借鉴决策理论中的结构化专家判断方法,利用上下文相关的校准问题,基于模型的概率预测质量评估其可靠性。具体采用Cooke风格的对数加权策略,惩罚过度自信的错误预测,奖励校准良好的专家。在MMLU和MMLU-Pro数据集上,对同质、异质及受污染专家集合进行评估。结果显示,在同质情况下各方法表现相近,但在异质与污染环境下,Cooke加权显著提升准确率-可靠性平衡,并保持对不可靠专家的鲁棒性。结论表明,多模型聚合不仅需融合预测,更需在不确定性下校准信任。

原文摘要 · Abstract (English)

Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs. However, existing aggregation methods typically assume that all models are equally trustworthy, overlooking differences in uncertainty quality. This assumption is poorly suited to heterogeneous LLMs, whose reliability and capability vary significantly, making naive aggregation vulnerable to unreliable or adversarial experts. In this work, we formulate multi-LLM aggregation as a problem of uncertainty-aware trust estimation. We adapt structured expert judgment from decision theory, using context-aware calibration questions to estimate expert reliability based on the quality of its probabilistic predictions. Specifically, we employ Cooke-style log weighting, which penalises overconfident incorrect predictions and favours well-calibrated experts. We evaluate our approach on MMLU and MMLU-Pro across homogeneous, heterogeneous, and contaminated expert panels. Results show that while aggregation methods perform similarly in homogeneous settings, Cooke weighting becomes critical under heterogeneity and contamination. It achieves a superior accuracy-reliability balance and remains robust when unreliable experts are introduced. These findings suggest that Multi-LLM aggregation requires not just combining predictions, but calibrating trust under uncertainty.

多模型集成可信度评估不确定性量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。