arXiv:2504.14154cs.CLcs.AI2025-04ACL被引 25

提出新方法精准识别大模型预测中的异常不确定样本

SConU: Selective Conformal Uncertainty in Large Language Models

  • 引入显著性检验与双重校准p值,检测偏离不确定性分布的异常样本
  • 在单领域与跨领域场景下实现可控误覆盖率,提升预测可靠性
  • 适合高风险问答任务,帮助用户理解模型置信度边界

随着大语言模型在现实应用中日益普及,确保特定任务指标的可靠性至关重要。现有基于分割共形预测的不确定性方法虽能提供用户指定的正确性覆盖率,但常无法识别违反交换性假设的异常样本,导致误覆盖率无界且预测集不可操作。本文提出首个实现显著性检验的新型方法SConU,通过构建两个共形p值,可在特定可管理风险水平下判断样本是否偏离校准集的不确定性分布。该方法不仅在单领域与跨领域场景中严格控制误覆盖率,还提升了预测效率。我们进一步分析共形过程各组件,旨在逼近条件覆盖,尤其适用于高风险问答任务。

原文摘要 · Abstract (English)

As large language models are increasingly utilized in real-world applications, guarantees of task-specific metrics are essential for their reliable deployment. Previous studies have introduced various criteria of conformal uncertainty grounded in split conformal prediction, which offer user-specified correctness coverage. However, existing frameworks often fail to identify uncertainty data outliers that violate the exchangeability assumption, leading to unbounded miscoverage rates and unactionable prediction sets. In this paper, we propose a novel approach termed Selective Conformal Uncertainty (SConU), which, for the first time, implements significance tests, by developing two conformal p-values that are instrumental in determining whether a given sample deviates from the uncertainty distribution of the calibration set at a specific manageable risk level. Our approach not only facilitates rigorous management of miscoverage rates across both single-domain and interdisciplinary contexts, but also enhances the efficiency of predictions. Furthermore, we comprehensively analyze the components of the conformal procedures, aiming to approximate conditional coverage, particularly in high-stakes question-answering tasks.

大模型不确定性共形预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。