提出新方法量化大模型上下文学习中的随机不确定性,提升预测可信度评估
Quantifying Aleatoric Uncertainty of In-Context Learning for Robust Measure of LLM Prediction Confidence

- 基于贝叶斯视角与模型内部表示,构建自函数向量以捕捉上下文学习中的潜在概念
- 在合成与真实数据上验证,可更准确分离并测量随机不确定性,优于现有方法
- 适用于幻觉检测等可信应用,推动不确定性量化与模型机制理解的结合
上下文学习(ICL)使大语言模型通过少量示例适应新任务,但其可靠性存疑:预测结果对提示设计和模型理解能力高度敏感,难以区分失败是源于数据特性还是模型局限。不确定性分解——将随机性与认知性不确定性分离——在此场景中尤为关键,但现有方法针对标准生成任务设计,无法捕捉ICL的独特动态。为此,我们引入自函数向量概念,基于贝叶斯视角与ICL的机械可解释性,利用模型内部表示建模上下文提示中学习到的潜在概念,从而在贝叶斯框架内直接估计随机不确定性,避免依赖脆弱的输入或解码操作。鉴于缺乏成熟基准与评估协议,我们还提出了首个严谨的评估框架,通过受控方式操控数据,精确量化随机不确定性并将其与认知不确定性分离。在合成任务中建立概念基础后,扩展至真实数据集,结果显示本方法比现有方法更可靠地衡量了ICL下的预测不确定性。此外,该方法可作为实用工具用于幻觉检测等可信应用。研究为不确定性量化与模型行为机制理解的结合开辟新路径。
原文摘要 · Abstract (English)
In-Context Learning (ICL) allows LLMs to adapt to new tasks from a few demonstrations, but its reliability remains a concern: predictions are highly sensitive to both prompt design and the model's ability to understand the context, obscuring whether failures arise from data properties or model limitations. Uncertainty decomposition-separating aleatoric from epistemic sources-is particularly crucial in this setting, yet existing methods, designed for standard generation tasks, fail to capture the unique dynamics of ICL. To address this, we introduce a concept of self-function vectors, built upon Bayesian views and the mechanistic interpretability of ICL. These vectors leverage internal model representations to model the latent concept learned during in-context prompting, thereby enabling a direct estimation of aleatoric uncertainty within a Bayesian framework and circumventing the reliance on brittle input or decoding manipulations. Given the lack of established benchmarks and suitable evaluation protocols, we also propose the first and rigorous evaluation protocol, in which data is manipulated in controlled ways so as to quantify aleatoric uncertainty precisely and separately from epistemic uncertainty. With this new evaluation framework, initially grounded in synthetic tasks for conceptual development and subsequently extended to real-world datasets, we show that our proposed methodology can measure uncertainty of LLM predictions made under ICL more reliably than existing alternative methods. Moreover, we show it can be used as a practical tool for trustworthy-related applications, such as hallucination detection. Our findings pave a new direction for connecting the quantitative view of uncertainty with the mechanistic understanding of model behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。