arXiv:2605.13484cs.LGcs.AI2026-05

发现大模型在不同输入上的校准偏差,定位错误高发区域。

Discovery of Hidden Miscalibration Regimes

论文配图:Discovery of Hidden Miscalibration Regimes
图 1 · 摘自论文原文
  • 通过学习输入空间的校准感知表示,识别局部校准异常。
  • 在12个大模型上发现输入依赖的校准异质性普遍存在。
  • 可定位校准失败区,提升局部置信度修正效果。

校准通常通过比较模型置信度与实际正确率来评估,隐含假设可靠性仅取决于置信度得分。然而,这种视角可能掩盖重要结构:模型在某些输入类型上系统性过自信,而在另一些上则过不自信,导致全局校准诊断掩盖了局部失效。为此,我们提出在无需预定义数据切片的前提下,发现隐藏的校准偏差区域。定义相应的校准场,并提出一种估计框架。方法在学习的输入空间几何中,通过核平滑估计有符号局部校准偏差。在四个真实世界的大语言模型基准和十二个大模型上,我们发现输入依赖的校准异质性广泛存在。进一步表明,所发现的校准场具有可操作性:支持局部置信度修正,在传统方法如等熵回归和温度缩放表现较弱的系统性校准偏差区域,显著降低校准误差。

原文摘要 · Abstract (English)

Calibration is commonly evaluated by comparing model confidence with its empirical correctness, implicitly treating reliability as a function of the confidence score alone. However, this view can hide substantial structure: models may be systematically overconfident on some kinds of inputs and underconfident on others, causing global reliability diagnostics to obscure localised calibration failures. To address this, we formulate the problem of discovering hidden miscalibration regimes without assuming access to predefined data slices. We define the corresponding miscalibration field and propose a diagnostic framework for estimating it. Our approach learns a calibration-aware representation of the input space and estimates signed local miscalibration by kernel smoothing in the learned geometry. Across four real-world LLM benchmarks and twelve LLMs, we find that input-dependent calibration heterogeneity is prevalent. We further show that the discovered fields are actionable: they support local confidence correction and reduce calibration error in systematically miscalibrated regions where confidence-based methods such as isotonic regression and temperature scaling are less effective.

大模型校准置信度修正异常检测模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。