提出MAPS算法,实现高维数据下无需假设的精准预测区间
The MAPS Algorithm: Fast model-agnostic and distribution-free prediction intervals for supervised learning
- 基于提升预测模型,用自助法构建分布无关的条件预测区间
- 在高维输入下仍有效,且能处理异方差误差,保证个体层面覆盖
- 适合需要可靠不确定性估计的模型应用,如图像分类与贝叶斯推断
现代监督学习中的核心挑战是在高维场景下生成可靠的条件预测区间:现有方法常依赖严格建模假设,无法随特征维度增长而扩展,或仅保证边际(总体)覆盖而非条件(个体)覆盖。本文提出新的条件表示——提升预测模型(LPM),并引入MAPS(模型无关预测集)算法,可为任意训练好的预测模型生成分布自由的条件预测区间。该方法基于自助法,具备高维可扩展性,并能捕捉异方差误差。我们建立了LPM的理论性质,揭示了预测精度与区间长度的关系,并给出渐近条件覆盖的充分条件。通过模拟研究评估了MAPS在有限样本下的表现,并应用于基于模拟的推断和图像分类任务。在前者中,MAPS首次实现神经贝叶斯估计器的去偏及参数置信区间的构造;在后者中,首次考虑了模型校准与标签预测中的不确定性。
原文摘要 · Abstract (English)
A fundamental problem in modern supervised learning is computing reliable conditional prediction intervals in high-dimensional settings: existing methods often rely on restrictive modelling assumptions, do not scale as predictor dimension increases, or only guarantee marginal (population-level) rather than conditional (individual-level) coverage. We introduce the $\textit{lifted predictive model}$ (LPM), a new conditional representation, and propose the MAPS (Model-Agnostic Prediction Sets) algorithm that produces distribution-free conditional prediction intervals and adapts to any trained predictive model. Our procedure is bootstrap-based, scales to high-dimensional inputs and accounts for heteroscedastic errors. We establish the theoretical properties of the LPM, connect prediction accuracy to interval length, and provide sufficient conditions for asymptotic conditional coverage. We evaluate the finite-sample performance of MAPS in a simulation study, and apply our method to simulation-based inference and image classification. In the former, MAPS provides the first approach for debiasing neural Bayes estimators and constructing valid confidence intervals for model parameters given the estimators, at any desired level. In the latter, it provides the first approach that accounts for uncertainty in model calibration and label prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。