提出权重裁剪方法,解决数据分布偏移下预测置信集失效问题。
Weight Clipping for Robust Conformal Inference under Unbounded Covariate Shifts

- 用裁剪最小二乘法估计密度比,降低方差
- 理论证明可实现受控的覆盖不足,且不随高阶矩膨胀
- 能自动估计需调整的置信水平,适合真实数据场景
分位数预测(CP)提供无需分布假设的预测集,但其保证依赖训练与测试数据的可交换性,而实际中常因协变量偏移破坏。加权分位数预测(WCP)虽可处理偏移,但在密度比无界或需学习时,易出现严重覆盖不足,源于密度比过拟合及非符合性得分阈值估计方差过高。为此,本文提出裁剪最小二乘重要性拟合(CLISF),一种低方差密度比估计方法。我们证明,将CLISF学习的密度比用于WCP时,预期覆盖不足有界。进一步,通过略提高覆盖目标可校正覆盖不足,且能从数据中估计所需膨胀量。这是首个关于权重裁剪在分位数推断中的理论保证,实现了数据集条件下的覆盖性,样本复杂度不随真实密度比的高阶矩增长——这是此前工作的关键局限。我们在真实世界基准和合成数据上验证了结果。
原文摘要 · Abstract (English)
Conformal prediction (CP) provides powerful, distribution-free prediction sets, but its guarantees rely on the exchangeability of training and test data, which is often violated in practice due to covariate shifts. While weighted conformal prediction (WCP) is designed to handle such shifts, it can suffer from significant undercoverage when the density ratio between the distributions is unbounded and/or must be learned. This is because of both overfitting in learning the density ratio, and high variance in estimating the nonconformity score threshold. To address this, we introduce clipped least-squares importance fitting (CLISF) as a reduced-variance method for density ratio estimation. Specifically, we show that density ratios learned using CLISF, when plugged into WCP, have bounded expected undercoverage. Furthermore, we show that the undercoverage can be corrected by running WCP with a slightly inflated coverage target; crucially, we are able to estimate the required level of inflation from the data. We provide the first theoretical guarantees for weight clipping in conformal inference, achieving dataset-conditional coverage with a sample complexity that does not blow up with the higher moments of the true density ratio -- a key limitation of prior work. We verify our results on real-world benchmarks and synthetic data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。