让大模型根据答案调整自信表达,解决过度自信问题。
ADVICE: Answer-Dependent Verbalized Confidence Estimation
- 通过微调使自信表达依赖模型自身答案
- 显著提升自信校准度,且不降低任务表现
- 适合关注模型可信度与可解释性的研究者
大语言模型(LLMs)能够以自然语言表达自信程度,提升了透明性与可靠性。然而,这种表达常伴随系统性过度自信,其根源尚不明确。本文分析了自信表达的动态机制,发现‘答案无关性’——即模型未能根据自身答案调整自信——是导致该问题的主要原因。为此,提出ADVICE(答案依赖的自信表达微调框架),通过强化答案相关的自信估计来改善这一现象。大量实验表明,ADVICE显著提升了自信校准效果,并在未见场景中表现出强泛化能力,同时不损害任务性能。进一步分析显示,这些改进源于增强的答案依赖性,揭示了过度自信的成因,为可信的自信表达提供了支持。
原文摘要 · Abstract (English)
Recent progress in large language models (LLMs) has enabled them to communicate their confidence in natural language, improving transparency and reliability. However, this expressiveness is often accompanied by systematic overconfidence, whose underlying causes remain poorly understood. In this work, we analyze the dynamics of verbalized confidence estimation and identify answer-independence -- the failure to condition confidence on the model's own answer -- as a primary driver of this behavior. To address this, we introduce ADVICE (Answer-Dependent Verbalized Confidence Estimation), a fine-tuning framework that promotes answer-grounded confidence estimation. Extensive experiments show that ADVICE substantially improves confidence calibration, while exhibiting strong generalization to unseen settings without degrading task performance. We further demonstrate that these gains stem from enhanced answer dependence, shedding light on the origins of overconfidence and enabling trustworthy confidence verbalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。