提出一种新方法,精准拆分大模型预测的两种不确定性。
Fine-Grained Uncertainty Decomposition in Large Language Models: A Spectral Approach
- 用量子信息中的冯诺依曼熵分解模型不确定性
- 在多个数据集上优于现有最先进方法
- 适合需要可信度评估的AI应用开发者
随着大语言模型(LLMs)在各类应用中广泛应用,获取其预测不确定性的可靠度量变得至关重要。准确区分源自输入数据内在模糊性的偶然不确定性(aleatoric uncertainty)与仅源于模型局限性的认知不确定性(epistemic uncertainty),是有效应对两类不确定性来源的关键。本文提出一种名为谱不确定性(Spectral Uncertainty)的新方法,用于量化和分解LLMs中的不确定性。该方法基于量子信息理论中的冯诺依曼熵,为将总不确定性分离为独立的偶然与认知成分提供了严格的理论基础。与现有基线方法不同,本方法引入细粒度的语义相似性表示,能够对模型输出中的多种语义解释进行更细致的区分。实证评估表明,谱不确定性在多种模型和基准数据集上均优于当前最先进的方法,在估计偶然不确定性与总不确定性方面表现更优。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are increasingly integrated in diverse applications, obtaining reliable measures of their predictive uncertainty has become critically important. A precise distinction between aleatoric uncertainty, arising from inherent ambiguities within input data, and epistemic uncertainty, originating exclusively from model limitations, is essential to effectively address each uncertainty source. In this paper, we introduce Spectral Uncertainty, a novel approach to quantifying and decomposing uncertainties in LLMs. Leveraging the Von Neumann entropy from quantum information theory, Spectral Uncertainty provides a rigorous theoretical foundation for separating total uncertainty into distinct aleatoric and epistemic components. Unlike existing baseline methods, our approach incorporates a fine-grained representation of semantic similarity, enabling nuanced differentiation among various semantic interpretations in model responses. Empirical evaluations demonstrate that Spectral Uncertainty outperforms state-of-the-art methods in estimating both aleatoric and total uncertainty across diverse models and benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。