评估大模型在问答任务中对认知与非认知不确定性的量化能力,揭示不同方法的适用场景。
Measuring Aleatoric and Epistemic Uncertainty in LLMs: Empirical Evaluation on ID and OOD QA Tasks
- 基于概率和语义一致性的方法分别适用于分布内与分布外场景
- 信息类方法在分布内表现优异,密度类方法在分布外更可靠
- 多指标验证下语义一致性方法具有跨数据集稳定性
大型语言模型(LLMs)在众多领域广泛应用,其输出可信度至关重要,不确定性估计(UE)在此扮演关键角色。本文针对分布内(ID)与分布外(OOD)问答任务,对十二种不同的不确定性估计方法进行系统性实证研究,采用包括来自大模型批评者(LLMScore)在内的四种生成质量评估指标。分析表明:基于信息的方法(利用标记与序列概率)在分布内表现突出,与其对数据的理解对齐;而基于密度的方法及P(True)指标在分布外场景中更具优势,能有效捕捉模型的认知不确定性;语义一致性方法通过评估生成答案的变异性,在各类数据集与评估指标下均表现稳健,虽非全局最优,但具备广泛适用性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have become increasingly pervasive, finding applications across many industries and disciplines. Ensuring the trustworthiness of LLM outputs is paramount, where Uncertainty Estimation (UE) plays a key role. In this work, a comprehensive empirical study is conducted to examine the robustness and effectiveness of diverse UE measures regarding aleatoric and epistemic uncertainty in LLMs. It involves twelve different UE methods and four generation quality metrics including LLMScore from LLM criticizers to evaluate the uncertainty of LLM-generated answers in Question-Answering (QA) tasks on both in-distribution (ID) and out-of-distribution (OOD) datasets. Our analysis reveals that information-based methods, which leverage token and sequence probabilities, perform exceptionally well in ID settings due to their alignment with the model's understanding of the data. Conversely, density-based methods and the P(True) metric exhibit superior performance in OOD contexts, highlighting their effectiveness in capturing the model's epistemic uncertainty. Semantic consistency methods, which assess variability in generated answers, show reliable performance across different datasets and generation metrics. These methods generally perform well but may not be optimal for every situation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。