arXiv:2509.25532cs.CLcs.AI2025-09被引 12

让大模型自己生成干扰项,更准确地评估自身信心。

Calibrating Verbalized Confidence with Self-Generated Distractors

  • 用模型自动生成干扰项,计算信心一致性来纠正过度自信
  • 在10次推理下表现优于自洽性方法100次,误差更低
  • 适合需要可靠置信度的AI安全与可信应用

大语言模型(LLM)输出的信心估计对用户信任至关重要。尽管模型能以人类可理解的方式表达信心,但实证发现其口头信心常失准——在低准确率任务上仍报高信心,损害信任与安全。我们假设这种过度自信源于模型对信息量少的陈述更具可诱导性,并实证验证了低准确率陈述上的可诱导性更高。基于此,我们提出Distractor-Normalized Coherence(DINCO),通过让模型在多个自生成的干扰项(即替代陈述)上独立表达信心,并以总信心值进行归一化,从而估计并校正可诱导性偏差。为进一步提升校准效果,我们引入生成器-验证器不一致机制,将归一化后的验证器信心与基于一致性的生成器信心结合。我们将自洽性视为多采样生成间的连贯性,而归一化口头信心则体现对不相容陈述验证间的连贯性,使两者互补融合于DINCO。分析表明,DINCO提供更非饱和、更可用的信心估计;单纯增加采样次数无法弥补与基线差距,且10次推理的DINCO性能超过自洽性在100次推理下的表现。

原文摘要 · Abstract (English)

Calibrated confidence estimates are necessary for large language model (LLM) outputs to be trusted by human users. While LLMs can express their confidence in human-interpretable ways, verbalized LLM-generated confidence scores have empirically been found to be miscalibrated, reporting high confidence on instances with low accuracy and thereby harming trust and safety. We hypothesize that this overconfidence often stems from a given LLM's heightened suggestibility when faced with claims that it encodes little information about; we empirically validate this hypothesis, finding more suggestibility on lower-accuracy claims. Building on this finding, we introduce Distractor-Normalized Coherence (DINCO), which estimates and accounts for an LLM's suggestibility bias by having the model verbalize its confidence independently across several self-generated distractors (i.e. alternative claims), and normalizes by the total verbalized confidence. To further improve calibration, we leverage generator-validator disagreement, augmenting normalized validator confidence with a consistency-based estimate of generator confidence. Here, we frame the popular approach of self-consistency as leveraging coherence across sampled generations, and normalized verbalized confidence as leveraging coherence across validations on incompatible claims, allowing us to integrate these complementary dimensions of coherence into DINCO. Moreover, our analysis shows that DINCO provides less saturated -- and therefore more usable -- confidence estimates, and that further sampling alone cannot close the gap between DINCO and baselines, with DINCO at 10 inference calls outperforming self-consistency at 100.

信心校准大模型可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。