arXiv:2608.13591cs.AIcs.CL2026-08中稿 · the 2nd Workshop o…

发现大模型高自信错误可能并非脆弱,而是稳定错误。

Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors

论文配图:Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors
图 1 · 摘自论文原文
  • 用输出级审计评分与内部敏感性探测结合分析模型错误稳定性。
  • 自省提示可降低多层隐藏状态敏感度,但不改善校准度。
  • 适合关注模型可信度与鲁棒性的研究者阅读。

大语言模型的高自信错误常被视为内部推理脆弱的证据。我们提出另一种可能:稳定误校准,即一个自信的错误答案在小扰动下仍保持局部稳定。通过两种诊断方法——基于标签的输出级审计评分(按置信度变化和过度自信错误排序)与内部敏感性探针(测量隐藏状态移动),我们在一个多领域二分类事实审计数据集上发现,该审计评分能追踪到弃权感知自省可减少决策损失的领域;而直接标注基线则更强烈地反映相同收益。内部分析显示,自省提示在三个开源模型中均一致降低各层隐藏状态敏感度。这支持了提示引发的局部稳定,而非仅输出层面的弃权模式,但不意味着校准:审计定义的过度自信错误并未比自信正确答案更敏感,因此部分高自信错误可能是稳定且误校准的,而非简单脆弱。

原文摘要 · Abstract (English)

High-confidence errors in large language models are often treated as evidence of fragile internal inference. We study a different possibility: stable miscalibration, where a confident wrong answer remains locally stable under small perturbations. We combine two diagnostics: a label-aware output-level audit score that ranks domains by confidence variation and overconfident mistakes under a forced-answer baseline, and an internal sensitivity probe that measures hidden-state movement. On a multi-domain binary factual audit set, this audit score tracks where abstention-aware self-critique reduces decision loss, although direct labeled baselines rank the same gain more strongly. Internally, self-critical prompting consistently reduces hidden-state sensitivity across layers in three open-weight models. This supports prompt-induced local stabilization rather than a purely output-level abstention pattern, but it does not imply calibration: audit-defined overconfident errors are not clearly more locally sensitive than confidently correct answers, so some high-confidence errors may be stable and miscalibrated rather than simply fragile.

大模型置信度错误分析自省提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。