通过词级别温度缩放提升语言模型问答的不确定性量化精度
Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling
- 引入词级别温度缩放,优化模型置信度分布
- 单标量温度调优显著改善校准度与区分度
- 适合关注模型可靠性与可信AI的研究者
校准是可靠语义不确定性量化的核心,但以往研究主要关注区分度,忽视了校准。由于校准与区分度反映不确定性不同方面,仅关注区分度会导致认知不完整。我们系统评估了多种置信度度量在两个维度的表现。结果表明,现有方法(尤其是固定温度启发式)产生系统性校准偏差且区分度差。我们证明,优化单一标量温度(我们认为这提供了合适的归纳偏置)是一种出人意料地简单而有效的方法。全面评估证实,温度缩放在问答任务中持续提升语义校准度、区分度和下游熵,优于启发式基线及更复杂的词级别重校准方法。
原文摘要 · Abstract (English)
Calibration is central to reliable semantic uncertainty quantification, yet prior work has largely focused on discrimination, neglecting calibration. As calibration and discrimination capture distinct aspects of uncertainty, focusing on discrimination alone yields an incomplete picture. We address this gap by systematically evaluating both aspects across a broad set of confidence measures. We show that current approaches, particularly fixed-temperature heuristics, produce systematically miscalibrated and poorly discriminative semantic confidence distributions. We demonstrate that optimising a single scalar temperature, which, we argue, provides a suitable inductive bias, is a surprisingly simple yet effective solution. Our exhaustive evaluation confirms that temperature scaling consistently improves semantic calibration, discrimination, and downstream entropy, outperforming both heuristic baselines and more expressive token-level recalibration methods on question-answering tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。