让大模型说的自信和心里的自信对得上,提升可信度。
Direct Confidence Alignment: Aligning Verbalized Confidence with Internal Confidence In Large Language Models
- 用直接偏好优化让模型口头自信与内部概率对齐
- 在多个模型上提升自信表达一致性,但效果因架构而异
- 适合关注模型可解释性与可靠性研究者
随着大语言模型应用日益广泛,其可信性与可靠性愈发重要。校准旨在提升模型自信与其回答正确性的匹配度。然而,模型内部基于词元概率的内在自信与口头表达的外在自信常不一致,导致不同校准方法结果不可靠。本文提出直接自信对齐(DCA),通过直接偏好优化使模型的口头自信与内部自信对齐,而非直接追求真实准确率,从而增强模型透明性与可靠性。我们在多个开源大模型上评估DCA,并引入三项基于校准误差的新指标。结果表明,DCA在部分模型架构上显著改善了自信对齐度,减少了自信表达不一致现象;但对其他模型效果有限,凸显出实现更可解释、可信大模型需发展更依赖模型特性的方法。
原文摘要 · Abstract (English)
Producing trustworthy and reliable Large Language Models (LLMs) has become increasingly important as their usage becomes more widespread. Calibration seeks to achieve this by improving the alignment between the model's confidence and the actual likelihood of its responses being correct or desirable. However, it has been observed that the internal confidence of a model, derived from token probabilities, is not well aligned with its verbalized confidence, leading to misleading results with different calibration methods. In this paper, we propose Direct Confidence Alignment (DCA), a method using Direct Preference Optimization to align an LLM's verbalized confidence with its internal confidence rather than ground-truth accuracy, enhancing model transparency and reliability by ensuring closer alignment between the two confidence measures. We evaluate DCA across multiple open-weight LLMs on a wide range of datasets. To further assess this alignment, we also introduce three new calibration error-based metrics. Our results show that DCA improves alignment metrics on certain model architectures, reducing inconsistencies in a model's confidence expression. However, we also show that it can be ineffective on others, highlighting the need for more model-aware approaches in the pursuit of more interpretable and trustworthy LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。