评估四种方法对大模型输出不确定性的估计效果,发现混合方法最可靠。
Systematic Evaluation of Uncertainty Estimation Methods in Large Language Models
- 对比四种不确定性估计方法,分析其在不同任务中的表现差异。
- 混合方法CoCoA在校准与区分正确答案上表现最佳,提升模型可靠性。
- 为实际应用中选择合适度量提供依据,适合关注模型可信度的开发者。
大型语言模型(LLMs)的输出具有不同的不确定性水平,且往往与正确性不一致,其实际可靠性难以保证。为量化这种不确定性,我们系统评估了四种用于估计LLM输出置信度的方法:VCE、MSP、Sample Consistency和CoCoA(Vashurin et al., 2025)。在使用一个先进的开源LLM进行的四项问答任务实验中,结果表明每种不确定性度量捕捉了模型置信度的不同方面,其中混合方法CoCoA整体表现最优,在校准性和正确答案区分能力上均有提升。我们讨论了各方法的权衡,并为实际应用中选择不确定性度量提供了建议。
原文摘要 · Abstract (English)
Large language models (LLMs) produce outputs with varying levels of uncertainty, and, just as often, varying levels of correctness; making their practical reliability far from guaranteed. To quantify this uncertainty, we systematically evaluate four approaches for confidence estimation in LLM outputs: VCE, MSP, Sample Consistency, and CoCoA (Vashurin et al., 2025). For the evaluation of the approaches, we conduct experiments on four question-answering tasks using a state-of-the-art open-source LLM. Our results show that each uncertainty metric captures a different facet of model confidence and that the hybrid CoCoA approach yields the best reliability overall, improving both calibration and discrimination of correct answers. We discuss the trade-offs of each method and provide recommendations for selecting uncertainty measures in LLM applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。