量化大模型推理时信心表达的可靠性,发现其常与真实信心不符。
Quantifying Faithful Confidence Expression in Large Reasoning Models

- 基于概率、隐状态和生成一致性,分析推理链中语言确定性。
- 多数大模型推理时信心表达不真实,提示优化无效。
- 适用于长链推理,可检测不同评估方法的脆弱性。
可靠的风险沟通对大语言模型的可信度至关重要,但模型内在信心与语言表达信心的一致性(即忠实校准,FC)始终存在缺陷。这一问题在大型推理模型(LRMs)中尤为突出,因其推理链常被用户视为深思熟虑、能力与自信的证据。尽管LRM广泛使用且重要性高,其信心表达是否忠实仍不明确。现有评估范式难以适应LRM产生的长链推理输出,因这些输出缺乏清晰步骤边界、结构不一致且蕴含复杂条件依赖,导致内在信心估计困难。为此,我们提出一种新框架,系统量化LRM的忠实校准能力。该框架基于三类内部不确定性源(词元概率、隐藏状态、采样一致性),分析语言上的确定性,并设计前缀控制采样法以消除不同推理路径间的条件与结构差异。在多种主流模型、数据集与提示下应用该框架,发现忠实信心表达仍是重大挑战:推理行为并不自动提升信心准确性,非推理模型的提示干预在推理场景中亦无效。不同信心估计算法对同一推理链给出截然不同的评估结果,暴露了既有评估方法的脆弱性。本工作确立了忠实校准作为LRM的独特可靠性与对齐目标,尤其在高风险场景日益普及的背景下更具意义。
原文摘要 · Abstract (English)
Reliable uncertainty communication is critical to the trustworthiness of LLMs, yet faithful calibration (FC)--the alignment between models' intrinsic and (linguistically) expressed confidence--is a persistent failure mode. This challenge is key for large reasoning models (LRMs), whose extended reasoning traces are often interpreted by users as evidence of deliberation, competence, and confidence. Despite the importance of FC and wide usage of LRMs, the extent to which LRMs can faithfully express their confidence remains poorly understood. Moreover, the prevailing paradigm to measure FC does not generalize well to the long chain-of-thought outputs generated by LRMs, which tend to lack clear step boundaries, involve inconsistent step structure, and encode complex conditional dependencies throughout the trace--complicating estimation of intrinsic confidence. To address this challenge, we introduce a novel framework to systematically quantify FC of LRMs. Our framework analyzes linguistic decisiveness relative to three sources of internal uncertainty, based on token probabilities, hidden states, and sampled response consistency. We also devise a prefix-conditioned sampling approach to control for conditional and structural variation across traces. Applying our framework to a diverse suite of leading models, datasets, and prompts, we find that faithful confidence expression is a significant challenge for LRMs. Reasoning behaviors do not automatically translate to improved FC, and prompt interventions for non-reasoning models do not improve faithfulness in the reasoning setting. Different confidence estimators further produce divergent assessments of the same traces, revealing fragility in prior evaluation methodologies. Taken together, our work establishes FC as a distinct reliability and alignment target for LRMs, particularly as such systems are increasingly deployed in high-stakes contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。