区分输入模糊性可显著提升大模型错误预测能力
The Role of Ambiguity in Error Prediction via Uncertainty Quantification

- 分离输入模糊性与不确定性信号,改进错误预测方法
- 在无歧义样本上,错误预测准确率提升超10点
- 适用于多种模型、数据集和不确定性来源
错误预测任务通常依赖不确定性量化(UQ)技术。然而,现有UQ指标不仅反映模型知识或能力不足,也包含输入本身固有的随机性(即似然不确定性)。本文提出一种新方法,通过解耦输入模糊性与UQ信号,提升大语言模型(LLM)的错误预测性能。我们在问答(QA)任务上使用六种UQ指标进行实验,发现这些指标在无歧义实例上的错误预测能力优于存在多个合理答案的情况。通过引入门控专家与选择性预测机制,将真实和预测的模糊性标签融入错误预测流程。结果表明,模糊性信息在不同模型家族、训练与评估范式、数据集(包括看似无歧义的数据集)及各类似然不确定性来源下均能提升预测效果,单个UQ指标在标准数据集上提升超过10点的PRR。
原文摘要 · Abstract (English)
The task of Error Prediction, namely predicting whether a model output is correct, is commonly tackled with Uncertainty Quantification (UQ). However, while uncertainty metrics capture when models lack knowledge or capacity to make a prediction, they also reflect aleatoric uncertainty, which is inherent in the model input and context. This paper presents a method for improving error prediction for Large Language Models (LLMs), by disentangling input ambiguity from UQ signal. We conduct experiments on the task of Question Answering (QA) with six UQ metrics and show that UQ metrics are more predictive of errors on unambiguous instances than on questions with multiple plausible answers. We use Gated Experts and Selective Prediction to incorporate gold and predicted ambiguity labels into the error prediction pipeline. We find that ambiguity information improves error prediction scores across model families, training and evaluation paradigms, datasets (including allegedly unambiguous ones), and sources of aleatoric uncertainty, yielding improvements of over 10 points of PRR for individual UQ metrics on standard datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。