提出新评估方法,让大模型的自信程度更靠谱地指导是否回答。
BAS: A Decision-Theoretic Approach to Evaluating Large Language Model Confidence

- 基于决策理论设计评分机制,衡量模型自信是否适合选择不答。
- 发现顶尖模型仍严重高估自己,易犯自信错误。
- 适合关注模型可靠性与安全决策的研究者和应用开发者。
大型语言模型(LLMs)在应答时常表现出过度自信但错误的答案,而标准评估方法要求必须回应,未考虑不同风险偏好下如何利用自信程度做出决策。为此,本文提出行为对齐得分(BAS),一种基于决策理论的评估指标,用于衡量大模型的自信程度在允许弃权情境下的决策支持能力。BAS源自明确的答题或弃权效用模型,通过整合多个风险阈值下的实际效用,得到依赖于自信程度大小与排序的决策可靠性度量。理论上证明:真实可信的自信估计能最大化期望BAS效用,将校准性与决策最优行为联系起来。BAS虽与逻辑损失等合理评分规则相关,但结构上存在差异——逻辑损失对低估和高估惩罚对称,而BAS则对高估错误施加不对称重罚。结合常用指标如ECE和AURC,我们构建了跨多个模型和任务的自报告自信可靠性基准。结果揭示决策可用自信存在显著差异,虽然更大更准确的模型通常拥有更高BAS,但前沿模型仍极易出现严重过自信。值得注意的是,具有相似ECE或AURC的模型可能表现截然不同的BAS,因过度自信错误所致,凸显传统指标局限性。此外,简单干预如前k个置信度提取和后验校准可显著提升自信可靠性。总体而言,本工作提供了一个原则性评估指标和全面基准,用于评测大模型的自信可靠性。
原文摘要 · Abstract (English)
Large language models (LLMs) often produce confident but incorrect answers in settings where abstention would be safer. Standard evaluation protocols, however, require a response and do not account for how confidence should guide decisions under different risk preferences. To address this gap, we introduce the Behavioral Alignment Score (BAS), a decision-theoretic metric for evaluating how well LLM confidence supports abstention-aware decision making. BAS is derived from an explicit answer-or-abstain utility model and aggregates realized utility across a continuum of risk thresholds, yielding a measure of decision-level reliability that depends on both the magnitude and ordering of confidence. We show theoretically that truthful confidence estimates uniquely maximize expected BAS utility, linking calibration to decision-optimal behavior. BAS is related to proper scoring rules such as log loss, but differs structurally: log loss penalizes underconfidence and overconfidence symmetrically, whereas BAS imposes an asymmetric penalty that strongly prioritizes avoiding overconfident errors. Using BAS alongside widely used metrics such as ECE and AURC, we then construct a benchmark of self-reported confidence reliability across multiple LLMs and tasks. Our results reveal substantial variation in decision-useful confidence, and while larger and more accurate models tend to achieve higher BAS, even frontier models remain prone to severe overconfidence. Importantly, models with similar ECE or AURC can exhibit very different BAS due to highly overconfident errors, highlighting limitations of standard metrics. We further show that simple interventions, such as top-$k$ confidence elicitation and post-hoc calibration, can meaningfully improve confidence reliability. Overall, our work provides both a principled metric and a comprehensive benchmark for evaluating LLM confidence reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。