arXiv:2608.08791cs.CL2026-08

发现扩散语言模型的置信度与真实表现严重不符,提出高效修复方案。

Unsure but Certain: Uncovering the Representation-Confidence Gap in Diffusion Language Models

  • 利用隐藏状态中的纠错信号改进答案排序
  • 噪声下标准模型反而更可靠,因置信度失真
  • 仅需轻量级工具,不改动原模型,无需重生成

扩散语言模型虽具较强上下文建模能力,但其内部对文本错误的检测准确率高,外部报告的置信度却无视该信号。当输入含噪时,模型准确率下降,置信度仍维持高位,答案排序能力退化至随机水平。这种内部表示与外部置信度之间的偏差称为‘表示-置信度鸿沟’。常见数学调整可消除置信度集中现象,但无法恢复排序性能。训练匹配可恢复准确率但不能修复排序缺陷;得分校准与输入级错误信号亦无法重排最终答案。然而,有效评估信息仍存在于隐藏状态中。我们提出一种轻量级提取工具,利用该信号提升排序,且完全冻结基模型、无需额外生成步骤。此方法证明了信号存在,但亦明确其局限。在噪声环境下,置信度可靠性比整体准确率更为关键。

原文摘要 · Abstract (English)

Diffusion language models use broad context to create text, suggesting they might handle input noise better than standard models. Testing reveals this is only partially true. Internally, diffusion models detect text errors highly accurately. Externally, their reported certainty ignores this signal. As accuracy drops due to noise, confidence stays near its maximum and the ability to correctly rank answers degrades toward random chance. We call this mismatch the representation confidence gap. The visible concentration of high certainty scores is a misleading surface symptom. Standard math adjustments remove this concentration but fail to fix the underlying loss of ranking order. This ranking deficit favors standard models under noisy conditions and resists common remedies. Matching training recovers accuracy but not ranking, while score recalibration and input level error signals cannot reorder the final answers. However, the information needed to properly evaluate an answer survives in the hidden states. A lightweight extraction tool uses this signal to improve ranking. This approach is highly efficient because it leaves the base model completely frozen and requires zero additional text generation steps. We present this tool to prove the signal exists, while clearly noting its limits. Ultimately, certainty reliability is a more pressing limit than overall accuracy under noisy conditions.

扩散模型置信度校准噪声鲁棒性语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。