arXiv:2607.19367cs.AI2026-07

提出新框架评估大模型置信度是否符合概率一致性,发现现有方法存在根本性缺陷。

Rethinking Uncertainty Evaluation in Large Language Models

论文配图:Rethinking Uncertainty Evaluation in Large Language Models
图 1 · 摘自论文原文
  • 从结构一致性、忠实性、实用性三方面定义可信置信度标准
  • 实测31%情况下模型对逻辑更简单的问题反而信心更低
  • 现有校准方法无法解决根本问题,适合关注模型可靠性研究者

校准是评估大模型置信度的主要标准,但存在明显不足:它允许平凡不一致的估计器,依赖评估分布,且无法检验估计是否可解释为一致的潜在概率函数。我们提出应使大模型置信度满足合理概率信念的条件,从结构一致性、忠实性和实用性三个维度构建C1评估指标。实证发现,广泛使用的估计器虽看似校准良好,却系统违反这些条件:模型在31%的情况下对逻辑更简单的问题给出更低信心;常见降低RMSCE的干预措施未改变结构性违规,表明校准与概率有效性无关。强化学习人类反馈(RLHF)和思维链(CoT)可提升实用性,但无法恢复一致性。结果表明当前大模型置信度无法被解释为一致概率,本框架提供了测量并弥合该差距的工具。

原文摘要 · Abstract (English)

Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and does not test the extent to which the estimation can be interpreted as a consistent, underlying probability function. What we actually need is for LLM confidence estimates to satisfy the conditions required of coherent probabilistic beliefs. We formalize these conditions along three axes (structural coherence, faithfulness, and usefulness) and operationalize them as the C1 metrics. Widely used estimators systematically violate these conditions despite appearing well-calibrated: models assign lower confidence to logically easier questions 31\% of the time, and common interventions reducing RMSCE leave structural violations unchanged, suggesting that calibration is orthogonal to probabilistic validity. RLHF and chain-of-thought improve usefulness metrics without restoring coherence. Our results show current LLM confidence estimates cannot be interpreted as coherent probabilities; our framework provides the tools to measure and close this gap.

大模型置信度概率校准可靠性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。