arXiv:2410.13735cs.LGstat.ME2024-10被引 8

用生成模型提升多维预测的不确定性量化精度

Generative Conformal Prediction with Vectorized Non-Conformity Scores

  • 用生成模型采样多个预测,构建向量化的非符合性分数
  • 在高置信区域给出更大预测集,低置信区更小或排除
  • 适合需要精准不确定性估计的多维预测场景

置信预测(CP)提供模型无关的不确定性量化并保证覆盖率,但传统方法在多维设置中常产生过于保守的不确定集,源于仅依赖预测误差的简单非符合性分数,未能捕捉误差分布的复杂性。为此,我们提出一种基于生成模型的置信预测框架,通过拟合数据分布采样多个预测,计算跨样本的非符合性分数,并在不同密度水平上估计经验分位数,构建基于密度排序的不确定性球。该方法实现更精确的不确定性分配——在高置信区域给出更大预测集,在低置信区域缩小或排除预测集,提升灵活性与效率。我们建立了统计有效性理论保证,并通过大量数值实验验证,该方法在合成和真实数据集上均优于现有最优技术。

原文摘要 · Abstract (English)

Conformal prediction (CP) provides model-agnostic uncertainty quantification with guaranteed coverage, but conventional methods often produce overly conservative uncertainty sets, especially in multi-dimensional settings. This limitation arises from simplistic non-conformity scores that rely solely on prediction error, failing to capture the prediction error distribution's complexity. To address this, we propose a generative conformal prediction framework with vectorized non-conformity scores, leveraging a generative model to sample multiple predictions from the fitted data distribution. By computing non-conformity scores across these samples and estimating empirical quantiles at different density levels, we construct adaptive uncertainty sets using density-ranked uncertainty balls. This approach enables more precise uncertainty allocation -- yielding larger prediction sets in high-confidence regions and smaller or excluded sets in low-confidence regions -- enhancing both flexibility and efficiency. We establish theoretical guarantees for statistical validity and demonstrate through extensive numerical experiments that our method outperforms state-of-the-art techniques on synthetic and real-world datasets.

不确定性量化生成模型置信预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。