忽略未观测序列会低估大模型不确定性,需纳入考虑
On the Role of Unobserved Sequences on Sample-based Uncertainty Quantification for LLMs
- 引入未观测输出序列的概率来改进不确定性估计
- 实验表明忽略未观测序列会导致不确定性被严重低估
- 适用于关注大模型安全与可信性的研究者
在大语言模型(LLMs)中量化不确定性对安全关键应用至关重要,有助于识别错误回答(即幻觉)。当前主流方法基于多次查询生成的输出序列及其概率,估算输出分布的熵。本文提出并实证表明,未观测序列的概率具有关键作用,建议未来研究将其纳入不确定性量化框架,以提升评估准确性。
原文摘要 · Abstract (English)
Quantifying uncertainty in large language models (LLMs) is important for safety-critical applications because it helps spot incorrect answers, known as hallucinations. One major trend of uncertainty quantification methods is based on estimating the entropy of the distribution of the LLM's potential output sequences. This estimation is based on a set of output sequences and associated probabilities obtained by querying the LLM several times. In this paper, we advocate and experimentally show that the probability of unobserved sequences plays a crucial role, and we recommend future research to integrate it to enhance such LLM uncertainty quantification methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。