改进语义字母表大小估计,提升大模型不确定性评估的准确性与可解释性。
Estimating Semantic Alphabet Size for LLM Uncertainty Quantification
- 提出改进的语义字母表大小估算方法,修正原始熵估计偏差。
- 仅用少量采样即实现更精准的不确定性量化,优于多个主流方法。
- 兼具高可解释性,适合需要透明决策的可信AI场景。
大量黑箱式大语言模型(LLM)不确定性量化方法依赖重复采样,计算成本高昂。因此,实际应用需从少量样本中可靠估计。语义熵(SE)是一种基于样本的不确定性估计器,其离散形式在黑箱设置下具有吸引力。近期对SE的扩展虽提升了幻觉检测能力,但引入了更多超参数,降低可解释性。本文重新审视经典的离散语义熵(DSE)估计器,发现其低估真实语义熵,符合理论预期。为此,我们提出一种改进的语义字母表大小估计方法,并证明将其用于调整DSE以考虑样本覆盖度后,能显著提升在目标场景下的语义熵估计精度。此外,我们验证了两种语义字母表大小估计器(包括新方法)在识别错误LLM响应方面表现不逊于多个顶尖方法,且保持高度可解释性。
原文摘要 · Abstract (English)
Many black-box techniques for quantifying the uncertainty of large language models (LLMs) rely on repeated LLM sampling, which can be computationally expensive. Therefore, practical applicability demands reliable estimation from few samples. Semantic entropy (SE) is a popular sample-based uncertainty estimator with a discrete formulation attractive for the black-box setting. Recent extensions of SE exhibit improved LLM hallucination detection, but do so with less interpretable methods that admit additional hyperparameters. For this reason, we revisit the canonical discrete semantic entropy (DSE) estimator, finding that it underestimates the "true" semantic entropy, as expected from theory. We propose a modified semantic alphabet size estimator, and illustrate that using it to adjust DSE for sample coverage results in more accurate SE estimation in our setting of interest. Furthermore, we find that two semantic alphabet size estimators, including our proposed, flag incorrect LLM responses as well or better than many top-performing alternatives, with the added benefit of remaining highly interpretable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。