针对多变量预测的误差相关性,提出更精准的置信集构造方法。
Semiparametric conformal prediction
- 用非参数藤蔓耦合模型建模多维误差的联合分布
- 在真实数据上实现精确覆盖率与高效估计
- 适合处理标签缺失、多变量相关等复杂场景
许多风险敏感应用需要对多个可能相关的目标变量构建校准良好的预测集合,此时预测算法可能产生相关误差。本文旨在构建考虑向量型非符合度分数联合相关结构的合规范式预测集。基于多元分位数与半参数统计的丰富文献,我们提出一种算法来估计得分的 $1-α$ 分位数,其中 $α$ 为用户指定的误覆盖率。特别地,我们灵活使用非参数藤蔓耦合模型估计得分的联合累积分布函数(CDF),并通过其影响函数提升分位数估计的渐近效率。藤蔓分解使得该方法可扩展至大量目标变量。实验表明,该方法不仅保证渐近精确覆盖,还在多种真实回归问题上实现理想覆盖率和竞争力的效率,包括校准集中存在随机缺失标签的情况。
原文摘要 · Abstract (English)
Many risk-sensitive applications require well-calibrated prediction sets over multiple, potentially correlated target variables, for which the prediction algorithm may report correlated errors. In this work, we aim to construct the conformal prediction set accounting for the joint correlation structure of the vector-valued non-conformity scores. Drawing from the rich literature on multivariate quantiles and semiparametric statistics, we propose an algorithm to estimate the $1-α$ quantile of the scores, where $α$ is the user-specified miscoverage rate. In particular, we flexibly estimate the joint cumulative distribution function (CDF) of the scores using nonparametric vine copulas and improve the asymptotic efficiency of the quantile estimate using its influence function. The vine decomposition allows our method to scale well to a large number of targets. As well as guaranteeing asymptotically exact coverage, our method yields desired coverage and competitive efficiency on a range of real-world regression problems, including those with missing-at-random labels in the calibration set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。