arXiv:2506.05497cs.LGcs.AI2025-06NeurIPS被引 9

为生成模型设计无需结构假设的不确定性量化方法,提升安全部署可靠性。

Conformal Prediction Beyond the Seen: A Missing Mass Perspective for Uncertainty Quantification in Generative Models

  • 基于查询黑箱模型构建预测集,突破传统依赖输出结构的限制。
  • 引入缺失质量估计与新查询策略,实现覆盖率、查询预算与信息量的最优平衡。
  • 适用于任意黑箱大语言模型,显著提升语言任务预测集的信息量。

不确定性量化(UQ)对生成式AI模型(如大语言模型)的安全部署至关重要,尤其在高风险场景中。经典共形预测(CP)方法局限于回归与分类任务,依赖几何距离或softmax分数,这些工具预设了结构化输出。本文提出一种仅通过有限查询黑箱生成模型来构建预测集的新范式,揭示了覆盖率、测试时查询预算与预测集信息量之间的新权衡。我们提出共形预测查询框架(CPQ),其核心在于两个原则:一是最优查询策略,由缺失质量的衰减速率决定;二是最优样本映射至预测集的方法,依赖缺失质量本身。我们发展了新的缺失质量估计器与Good Turing估计算法。在三个真实开放任务及两种大语言模型上的精细实验表明,该方法可应用于任意黑箱语言模型,并验证了两原则对性能的独立贡献,且相比现有方法显著提升了语言不确定性量化下的预测集信息量。

原文摘要 · Abstract (English)

Uncertainty quantification (UQ) is essential for safe deployment of generative AI models such as large language models (LLMs), especially in high stakes applications. Conformal prediction (CP) offers a principled uncertainty quantification framework, but classical methods focus on regression and classification, relying on geometric distances or softmax scores: tools that presuppose structured outputs. We depart from this paradigm by studying CP in a query only setting, where prediction sets must be constructed solely from finite queries to a black box generative model, introducing a new trade off between coverage, test time query budget, and informativeness. We introduce Conformal Prediction with Query Oracle (CPQ), a framework characterizing the optimal interplay between these objectives. Our finite sample algorithm is built on two core principles: one governs the optimal query policy, and the other defines the optimal mapping from queried samples to prediction sets. Remarkably, both are rooted in the classical missing mass problem in statistics. Specifically, the optimal query policy depends on the rate of decay, or the derivative, of the missing mass, for which we develop a novel estimator. Meanwhile, the optimal mapping hinges on the missing mass itself, which we estimate using Good Turing estimators. We then turn our focus to implementing our method for language models, where outputs are vast, variable, and often under specified. Fine grained experiments on three real world open ended tasks and two LLMs, show CPQ applicability to any black box LLM and highlight: (1) individual contribution of each principle to CPQ performance, and (2) CPQ ability to yield significantly more informative prediction sets than existing conformal methods for language uncertainty quantification.

不确定性量化共形预测大语言模型缺失质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。