arXiv:2505.07459cs.IR2025-05ACL被引 23

现有不确定性估计方法在RAG中失效,研究提出新框架提升可靠性。

Why Uncertainty Estimation Methods Fall Short in RAG: An Axiomatic Analysis

  • 构建五条公理约束,评估不确定性的合理性
  • 实验表明现有方法均不满足全部公理,性能受限
  • 提出校准函数,提升估计与正确性相关性

大型语言模型(LLMs)虽在各类任务中表现优异,但也会产生错误或误导性输出。不确定性估计(UE)可量化模型置信度,帮助用户判断回答可靠性。然而,现有UE方法尚未在检索增强生成(RAG)场景中得到充分检验,该场景的提示词包含非参数化知识。本文表明,当前UE方法无法在RAG设置中可靠评估答案正确性。我们提出一个公理化框架,识别现有方法的缺陷并指导改进。该框架定义了五个约束条件,要求有效UE方法在引入检索文档后仍需满足。实验结果表明,无一现有方法完全满足所有公理,解释了其在RAG中的次优表现。我们进一步提出一种基于该框架的简单而有效的校准函数,不仅满足更多公理,且提升了不确定性估计与正确性之间的相关性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are valued for their strong performance across various tasks, but they also produce inaccurate or misleading outputs. Uncertainty Estimation (UE) quantifies the model's confidence and helps users assess response reliability. However, existing UE methods have not been thoroughly examined in scenarios like Retrieval-Augmented Generation (RAG), where the input prompt includes non-parametric knowledge. This paper shows that current UE methods cannot reliably assess correctness in the RAG setting. We further propose an axiomatic framework to identify deficiencies in existing methods and guide the development of improved approaches. Our framework introduces five constraints that an effective UE method should meet after incorporating retrieved documents into the LLM's prompt. Experimental results reveal that no existing UE method fully satisfies all the axioms, explaining their suboptimal performance in RAG. We further introduce a simple yet effective calibration function based on our framework, which not only satisfies more axioms than baseline methods but also improves the correlation between uncertainty estimates and correctness.

不确定性估计RAGLLM可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。