让大模型错得不自信,对的仍保持自信
Calibrating Overconfidence Without Sacrificing Confidence: Probe-Conditioned Head Intervention for LLMs

- 用冻结探针检测错误高自信回答,条件性调节注意力头输出
- 在OpenMathInstruct上将82.2%错误自信转为否定,降低ECE至9.2%
- 只损伤5.1%正确回答,适合需精准置信度的场景
大语言模型常对错误答案表现出过高信心。传统校准方法多全局或在分数层面操作,虽能降低不当自信,却可能削弱正确答案应有的信心。我们提出推理阶段的探针条件头干预(PCHI),利用冻结探针识别可能错误但自信的回答,并在生成信心时条件性地重标下游注意力头输出。在Qwen3-4B-Instruct解决OpenMathInstruct问题、使用结构化二元信心字段的任务中,读出标记的PCHI将82.2%原本错误的'yes'信心读出转为'no';跨上游信心模板标记的联合干预使ECE从21.9%降至9.2%,仅损害5.1%原本正确的'yes'读出。读出标记效应在Gemma3-4B上也显现,但上游干预效果较弱且更依赖掩码。结果表明,通过条件性内部干预可选择性降低口语化过度自信,部分解耦了抑制不当自信与保留正当自信之间的冲突。
原文摘要 · Abstract (English)
Large language models often express high confidence in answers that are wrong. Standard calibration remedies typically act globally or at the score level, reducing unwarranted confidence but also risking erosion of warranted confidence on correct answers. We introduce Probe-Conditioned Head Intervention (PCHI), an inference-time method that uses a frozen probe to detect likely wrong-but-confident responses and conditionally rescales downstream attention-head outputs during confidence generation. On Qwen3-4B-Instruct solving OpenMathInstruct problems with a structured binary confidence field, readout-token PCHI converts 82.2% of originally wrong-yes confidence readouts to $\texttt{no}$, while a joint intervention across upstream confidence-template tokens reduces ECE from 21.9% to 9.2% and damages only 5.1% of originally correct-yes readouts. The readout-token effect also appears on Gemma3-4B, though upstream interventions are weaker and more mask-dependent. These results show that verbalized overconfidence can be selectively reduced through conditionally applied internal intervention, partially decoupling the suppression of unwarranted confidence from the loss of warranted confidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。