提出新评估框架,检验大模型置信度在语言变化下的稳定性。
Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations
- 设计三类互补指标:抗扰动、语义等价答案稳定、语义差异敏感
- 多数现有方法在语义差异检测上表现差,难以区分错误与正确答案
- 适合关注模型可信度评估的开发者和应用部署者
置信度估计(CE)反映大语言模型回答的可靠性,影响用户信任与决策。现有评估多关注置信度与正确性的对齐,却忽略了语言变化的影响:置信度应在语义等价的提示或答案间保持一致,而在语义不同时变化,这可能反映正确性变化。为此,我们提出基于三个互补属性的新评估框架:对提示扰动的鲁棒性、对语义等价答案的稳定性、对语义差异答案的敏感性。实验表明这些指标与现有CE指标高度独立,且多数常用方法在语义差异检测上表现不佳,可能因未能有效利用生成端信息。整体框架揭示了当前置信度评估的盲区,并为真实场景中置信度估计器的选择提供指导。
原文摘要 · Abstract (English)
Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing evaluations mainly concern the alignment between confidence and correctness, but ignore the variability of language: confidence estimates should remain consistent under semantically equivalent prompts or answer variations, while changing when answer meaning differs, as this may indicate a change in correctness. Therefore, we introduce a novel evaluation framework based on three complementary properties: \textbf{robustness} to prompt perturbations, \textbf{stability} across semantically equivalent answers, and \textbf{sensitivity} to semantically different answers. We show that these metrics are largely independent from existing CE metrics, and that common CE methods often fail on them: while most methods achieve high robustness and stability, they struggle to distinguish semantically different answers, potentially because they do not effectively leverage generation-side information. Overall, our framework exposes overlooked limitations of current CE evaluations and provides guidance for selecting confidence estimators for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。