提出新方法分解大模型上下文学习中的不确定性,提升预测可靠性。
Variational Uncertainty Decomposition for In-Context Learning

- 用辅助查询作为探针,间接估计不确定性上界。
- 实验证明分解出的不确定性能准确反映认知与固有差异。
- 适合关注模型可信度与推理稳定性的研究者使用。
随着大语言模型在上下文学习中广泛应用,理解其预测中的不确定性来源变得至关重要。近期假设上下文学习可视为预测性贝叶斯推断,为贝叶斯不确定性估计开辟了路径,尤其支持将不确定性分解为因上下文数据不足导致的认知不确定性(epistemic)和任务本身固有的随机不确定性(aleatoric)。然而,由于底层贝叶斯模型中潜在参数后验难以计算,该分解思路尚未充分探索。本文提出一种变分不确定性分解框架,无需显式采样潜变量后验,通过优化辅助查询作为探针,获得模型上下文学习过程中异质性不确定性的上界,同时诱导出认知不确定性的下界。在合成与真实任务上的实验表明,该方法定量和定性地展示了分解出的不确定性具备理想的认知与固有不确定性特性。
原文摘要 · Abstract (English)
As large language models (LLMs) gain popularity in conducting prediction tasks in-context, understanding the sources of uncertainty in in-context learning becomes essential to ensuring reliability. The recent hypothesis of in-context learning performing predictive Bayesian inference opens the avenue for Bayesian uncertainty estimation, particularly for decomposing uncertainty into epistemic uncertainty due to lack of in-context data and aleatoric uncertainty inherent in the in-context prediction task. However, the decomposition idea remains under-explored due to the intractability of the latent parameter posterior from the underlying Bayesian model. In this work, we introduce a variational uncertainty decomposition framework for in-context learning without explicitly sampling from the latent parameter posterior, by optimising auxiliary queries as probes to obtain an upper bound to the aleatoric uncertainty of an LLM's in-context learning procedure, which also induces a lower bound to the epistemic uncertainty. Through experiments on synthetic and real-world tasks, we show quantitatively and qualitatively that the decomposed uncertainties obtained from our method exhibit desirable properties of epistemic and aleatoric uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。