大模型说的“有把握”和实际信心不一致,需多维度检验。
When Linguistic and Internal Confidence Diverge in Large Language Models
- 用逻辑值、语义熵等多指标对比模型内外信心
- 多数情况下语言自信与内部信心关联弱,越强模型越准
- 适合评估模型可靠性时做多轴诊断,别轻信口头自信
用户常要求大语言模型报告其置信度,但这种语言表达的自信是否反映模型内部真实信心尚不明确。本文在8个分类任务、2个生成任务和30个来自三个系列的模型上研究该问题。分类任务中,从关联性、数值一致性、校准性三个维度比较语言自信与基于logits的内部自信;生成任务中检验语言自信是否匹配语义熵衡量的不确定性。结果显示三者常出现分歧:实例级关联性平均较弱,但在更简单样本和更强基线模型上有所提升。指令微调模型通常报告更高自信,有时关联性更好,但存在更大信心差距且校准更差。提示设计主要改变报告自信的分布形态。态度线索会虚增自信但不提升对齐效果,而评分示例若避免置信值坍缩,可保留排序信号。回归分析表明,置信度分数的分布特性解释了大部分观察到的对齐模式,模型元数据的作用在控制后较小。这些结果支持语言自信为‘有损信道’的观点:更分散的语言自信分布可能携带有用排序信息,但无法实现校准。因此,在下游可靠性流程中使用语言自信前,应进行多轴诊断。
原文摘要 · Abstract (English)
Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's internal confidence. We study this question across 8 classification tasks, 2 generation tasks and 30 models from three families. For classification, we compare linguistic confidence with logits-based confidence along three axes: association, magnitude agreement and calibration. For generation, we test whether linguistic confidence tracks semantic-entropy-based uncertainty. The axes frequently diverge. Instance-level association is weak on average, although it improves on easier items and for stronger base models. Instruction-tuned models often report higher confidence and sometimes show higher association, but they also have larger confidence gaps and worse calibration. Prompt design mostly changes the distribution of reported confidence. Attitude cues inflate confidence without improving alignment, while score exemplars can preserve rank-order signal when they avoid collapsed confidence values. Regression analyses show that distributional properties of confidence scores explain much of the observed alignment pattern, with model metadata playing a smaller role after controls. These results support a lossy-channel view of linguistic confidence. A more dispersed verbal confidence distribution can carry useful rank information, but it does not make the scores calibrated. Linguistic confidence should therefore be evaluated with multi-axis diagnostics before being used in downstream reliability pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。