系统对比LLM不确定性度量与缓解方法,提供首个全面基准评测。
Comparing Uncertainty Measurement and Mitigation Methods for Large Language Models: A Systematic Review
- 构建针对LLM的不确定性量化与校准方法综合评测框架。
- 在三个主流可靠性数据集上验证七种方法,揭示性能差异。
- 适合关注大模型可信性、幻觉问题的研究者与工程师。
大型语言模型(LLMs)在多个领域具有变革性影响,但幻觉问题——即自信地输出错误信息——仍是其主要挑战之一。这引出关键问题:如何准确评估和量化LLM的不确定性?传统模型研究已探索不确定性量化(UQ)及校准技术以弥合不确定性与准确性之间的偏差。尽管部分方法已被适配至LLM,现有文献仍缺乏对其有效性的深入分析,也未提供可进行有意义比较的综合性基准。本文通过系统综述代表性工作并引入严谨的评测基准,填补该空白。基于三个广泛使用的可靠性数据集,我们实证评估了七种相关方法,验证了综述中的重要发现。最后,我们展望关键未来方向并梳理开放挑战。据我们所知,这是首个专门针对LLM校准方法与相关度量指标的系统性研究。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have been transformative across many domains. However, hallucination, i.e., confidently outputting incorrect information, remains one of the leading challenges for LLMs. This raises the question of how to accurately assess and quantify the uncertainty of LLMs. Extensive literature on traditional models has explored Uncertainty Quantification (UQ) to measure uncertainty and employed calibration techniques to address the misalignment between uncertainty and accuracy. While some of these methods have been adapted for LLMs, the literature lacks an in-depth analysis of their effectiveness and does not offer a comprehensive benchmark to enable insightful comparison among existing solutions. In this work, we fill this gap via a systematic survey of representative prior works on UQ and calibration for LLMs and introduce a rigorous benchmark. Using three widely used reliability datasets, we empirically evaluate seven related methods, which justify the significant findings of our review. Finally, we provide outlooks for key future directions and outline open challenges. To the best of our knowledge, this survey is one of the first dedicated studies to review the calibration methods and relevant metrics for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。