评估大模型黑箱不确定性估计方法,发现综合多信号的混合方法更可靠。
A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models

- 按思路将黑箱方法分为五类:语义化、采样、解释、多智能体和混合。
- 24种方法在4个模型4个数据集上测试,无单一方法全面领先。
- 在答案空间推理对比的方案表现更好,混合信号方法适应性强。
尽管大语言模型在众多任务中表现出强大能力,其输出仍可能不可靠且存在幻觉,因此不确定性估计对构建可信模型至关重要。实践中,主流大模型仅通过受限API访问,内部信号如logits和隐藏状态不可用,使得黑箱不确定性估计尤为关键。然而,现有黑箱不确定性估计研究方法零散,缺乏统一实证比较。为此,我们系统性地回顾了黑箱方法,并将其归纳为五类:基于语义化、采样、解释、多智能体及混合方法。我们构建统一评估框架,在4个模型和4个数据集设置下基准测试了24种代表性方法。结果表明,无单一方法在所有场景中持续领先。但能在答案空间进行推理与对比的方法通常表现更优,融合多种不确定信号的混合方法在多数条件下表现良好。通过发布基准数据和统一评估框架,我们旨在促进可复现比较,支持未来研究,同时为开发下一代黑箱不确定性估计方法提供实践指导。
原文摘要 · Abstract (English)
Although large language models (LLMs) have shown strong capabilities across a wide range of tasks, their outputs often remain unreliable and may contain hallucinations, making uncertainty estimation (UE) essential for building trustworthy LLMs. In practice, many mainstream LLMs are only accessible through restricted APIs, where internal signals such as logits and hidden states are unavailable, making black-box UE especially important. However, existing work on black-box UE for LLMs remains fragmented in methodology and lacks a unified empirical comparison. To address this gap, we present a systematic review of black-box UE methods and organize them into five categories: verbalization-based, sampling-based, explanation-based, multi-agent, and hybrid methods. We further build a unified evaluation framework and benchmark 24 representative methods across 4 models and 4 dataset settings. Our results show that no single method consistently dominates across all settings. Nevertheless, methods that reason over and compare candidates in the answer space are generally effective, and hybrid methods that combine multiple uncertainty signals perform well under most conditions. By releasing the benchmark data and a unified evaluation framework, we aim to facilitate reproducible comparisons and support future research, while our empirical findings provide practical guidance for developing future black-box UE methods for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。