首个大模型长文本问答不确定性校准基准,揭示现有方法的可靠性缺陷。
Benchmarking Uncertainty Calibration in Large Language Model Long-Form Question Answering
- 构建覆盖20个模型、68.5万条回答的科学问答评估框架
- 发现指令微调使词级别置信度失真,答案频率更可靠
- 批判仅用ECE评估校准,提醒方法选择需多维验证
大型语言模型广泛用于问答任务,尤其在自然科学领域。可靠的不确定性量化(UQ)对生成答案的可信采纳至关重要。现有UQ方法在依赖事实检索与推理的科学问答中验证不足。本文提出首个针对高推理需求问答的不确定性校准大规模基准,提供可复现的开源评估框架。研究涵盖20个基础、指令微调及推理增强型大模型,覆盖7个科学问答数据集(含多选与算术题),通过提示工程模拟开放问答场景。评估共68.5万条长文本回答,覆盖不同推理复杂度。在词级别,发现指令微调导致置信度分布极端化,降低其作为不确定性估计的可靠性;进一步推理微调虽加剧此现象,但推理过程可部分缓解,具体效果取决于模型提供商。在序列级别,显式表述方法系统性偏差明显且与正确性相关性弱,而答案频率(多次采样一致性)表现最优。研究还揭示仅依赖ECE作为评估指标会误导对UQ方法性能的判断。结果暴露当前大模型不确定性量化方法及评估实践的关键局限。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are commonly used in Question Answering (QA) settings, increasingly in the natural sciences if not science at large. Reliable Uncertainty Quantification (UQ) is critical for the trustworthy uptake of generated answers. Existing UQ approaches remain weakly validated in scientific QA, a domain relying on fact-retrieval and reasoning capabilities. We introduce the first large-scale benchmark for evaluating UQ metrics in reasoning-demanding QA studying calibration of UQ methods, providing an extensible open-source framework to reproducibly assess calibration. Our study spans up to 20 large language models of base, instruction-tuned and reasoning variants. Our analysis covers seven scientific QA datasets, including both multiple-choice and arithmetic question answering tasks, using prompting to emulate an open question answering setting. We evaluate and compare methods representative of prominent approaches on a total of 685,000 long-form responses, spanning different reasoning complexities representative of domain-specific tasks. At the token level, we find that instruction tuning induces strong probability mass polarization, reducing the reliability of token-level confidences as estimates of uncertainty. Models further fine-tuned for reasoning are exposed to the same effect, but the reasoning process appears to mitigate it depending on the provider. At the sequence level, we show that verbalized approaches are systematically biased and poorly correlated with correctness, while answer frequency (consistency across samples) yields the most reliable calibration. In the wake of our analysis, we study and report the misleading effect of relying exclusively on ECE as a sole measure for judging performance of UQ methods on benchmark datasets. Our findings expose critical limitations of current UQ methods for LLMs and standard practices in benchmarking thereof.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。