arXiv:2605.27752cs.AI2026-05被引 2

LLM自信度评估受提问方式影响,同一答案不同表述时可信度差异显著。

Same Answer, Different Confidence: Protocol Sensitivity in LLM Confidence Calibration

论文配图:Same Answer, Different Confidence: Protocol Sensitivity in LLM Confidence Calibration
图 1 · 摘自论文原文
  • 固定答案和正确性标签,对比模型自答与置信度提示下的得分表现。
  • 在12组测试中,9组下口语化自信度优于词元似然,且结果不因缩放改变。
  • 答案表达形式(如别名或标准写法)会影响自信度评分,需明确评估协议。

口头自信度是否比词元似然更校准?答案取决于如何测量词元似然:哪个答案被评分,以及在何种提示下。已有比较在此问题上存在分歧,且在十二项审计研究中,五项从未说明具体选择。我们为每个问题固定一次预测事件——模型自身的回答及其正确性标签,并在同一答案和标签下,分别在普通查询和置信度提示中进行评分。在四个问答数据集和三个7-8B Instruct模型上,这一设定改变了信号表现优劣的点估计,在ECE指标下有4组变化,在AUROC下有9组变化。由于AUROC对保持顺序的变换不变,排序差异无法由缩放解释。另外两种操作——用参考字符串替代模型答案,或读取首个词元而非完整答案——也表现出类似行为。跨三个答案槽位、两个评分字符串和两个读出方式,共产生十二种操作变体,在6组中使ECE比较结果不确定;而其他校准估计器影响较小,仅一个改变了单个提示上下文的胜者。此外,口语化自信度对答案表述敏感:将接受的TriviaQA别名替换为标准参考答案,自信度提升0.072,尽管两者均正确。因此,比较两种信号必须明确指定答案、上下文和评估协议。

原文摘要 · Abstract (English)

Is verbalized confidence better calibrated than token likelihood? The answer depends on how the token likelihood is measured: which answer is scored, and under which prompt. Published comparisons diverge on this, and in a twelve-study audit five never state the choice. We fix one prediction event per question, the model's own answer together with its correctness label, and score that same answer under a plain query and inside the confidence prompt, holding the answer and its label fixed. Across four QA datasets and three 7-8B Instruct models this changes which signal performs better, by point estimate, in 4 of 12 settings under ECE and 9 of 12 under AUROC. The AUROC result cannot come from rescaling the likelihoods, since AUROC is invariant to any common order-preserving transformation; the items are ordered differently. Two further choices behave the same way: substituting the reference string for the model's own answer, and reading the first answer token instead of the answer span. Crossing three answer slots, two scored strings, and two readouts gives twelve measured operational variants that leave the sign of the ECE comparison ndetermined in 6 of 12 settings, whereas alternative calibration estimators move it substantially less, although one changes a single prompted-context winner. Verbalized confidence is sensitive to answer formulation as well: replacing an accepted TriviaQA alias with the canonical reference raises confidence by $0.072$ although both answers are correct. Comparing the two signals therefore requires an explicit answer, context, and evaluation protocol.

大模型自信度校准提示工程评估协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。