用不确定性量化提升视觉语言模型的预测置信度,让结果更稳定可靠。
Empirical Bayes Conformal Prediction for Vision and Language Models

- 基于经验贝叶斯思想,将多个预测得分的波动性转化为新的置信度指标
- 在保持理论覆盖率的前提下,显著减少不稳定的错误预测被纳入结果集
- 适合对模型输出可靠性要求高的应用场景,如医疗、自动驾驶
分位数预测(CP)为现代视觉与语言模型提供无分布覆盖保证,但常因单个不稳定非符合性分数被迫做出排序判断。标准CP仅使用一次实现,而平均后校准方法虽通过多次实现平滑成点估计,却忽略了可变性信息——这种可变性本可用于识别候选是否真正稳定。弱信号可能因某次后验采样或提示措辞偶然表现良好而被纳入预测集。但波动性有助于区分真实信号与噪声波动。本文提出一种经验贝叶斯分位数预测框架,利用r值将得分变异性转化为带不确定性的非符合性分数。r值估计在考虑均值和不确定性后,某候选得分属于前排组的概率。该方法支持正态-正态闭式估计器与非参数后验采样估计器。以r值作为非符合性分数,在保持目标覆盖率的同时,于温和正则条件下严格减少高方差假阳性候选的纳入。在图像分类、基于CLIP的视觉语言模型基准以及大语言模型上,我们验证:当变异性具信息量时,r值分位数预测能维持覆盖率、提升排序稳定性并缩小预测集;当变异性消失时,则退化为传统CP行为。
原文摘要 · Abstract (English)
Conformal prediction (CP) gives distribution-free coverage for modern vision and language models, but it is often forced to make a ranking decision from a single unstable nonconformity score. Standard CP uses one realization, while average-then-calibrate variants smooth multiple realizations into a point estimate. Both options discard the inconsistency that can help identify whether a candidate is indeed stable. A weak answer can enter the conformal set even if the evidence is not strong, simply because one posterior sample or prompt phrasing made it look strong. But variability can help distinguish a stable signal from noise-driven fluctuations. We describe an empirical Bayes conformal prediction framework that uses $r$-values to convert score variability into an uncertainty informed nonconformity score. The resulting $r$-value estimates how likely a candidate's latent score belongs to the top-ranked group after accounting for both its mean score and its uncertainty. It admits both a closed-form Normal-Normal empirical Bayes estimator and a nonparametric posterior-sampling estimator. Using the $r$-value as the nonconformity score preserves the target conformal coverage while provably reducing the inclusion of high variance false candidates under mild regularity conditions. Across image classification, CLIP-based VLM benchmarks, and LLMs, we show that $r$-value conformal prediction preserves target coverage while improving ranking stability and reducing set size when variability is informative, and reverting to CP-like behavior when variability vanishes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。