提出一种新型分位数区间方法,提升预测区间在有限样本下的准确性与效率。
Conformalized Percentile Interval: Finite Sample Validity and Improved Conditional Performance

- 在概率积分变换空间中校准,利用神经网络估计的条件累积分布函数构造区间
- 实验显示区间长度显著缩短,且条件校准性能优于现有方法
- 适合需要高精度置信区间的实际场景,如金融、医疗预测
分位数预测提供了无需分布假设的有限样本边际覆盖预测区间,但在异方差、偏态响应或估计误差等复杂场景下,实现条件有效性与区间效率(如短区间长度)仍具挑战。本文提出一种基于概率积分变换(PIT)的分位数校准方法,利用神经网络估计的条件累积分布函数(CDF)构建有限样本调整后的分位数区间,其长度由估计的条件 CDF 确定。在 PIT 空间校准有效,因为当 CDF 估计准确时,PIT 值渐近特征无关,可缓解特征依赖性导致的误覆盖,改善条件校准。同时,分位数校准适应经验 PIT 分布,对条件 CDF 估计不完美具有鲁棒性。理论证明了所提方法的有限样本边际覆盖性,并在温和一致性条件下给出了渐近条件覆盖性。在多种合成与真实世界基准上的实验表明,该方法在条件校准和区间长度方面均显著优于现有方法。
原文摘要 · Abstract (English)
Conformal prediction provides distribution-free predictive intervals with finite-sample marginal coverage. However, achieving conditional validity and interval efficiency (in terms of short interval length) remains challenging, particularly in complex settings with heteroskedasticity, skewed responses, or estimation errors. We propose a conformal-style calibration method for responses obtained by the probability integral transform (PIT) of the conditional cumulative distribution function (CDF) estimated via neural networks to construct a finite-sample-adjusted percentile interval with the shortest length determined by the estimated conditional CDF. Calibrating in PIT space is effective because PIT values are asymptotically feature-independent when the CDF estimator is accurate, which mitigates feature-dependent miscoverage and improves conditional calibration. On the other hand, our percentile calibration adapts to the empirical PIT distribution, which is robust against a possibly imperfect estimation of the conditional CDF. We prove the finite-sample marginal coverage property of the proposed method and show its asymptotic conditional coverage under mild consistency conditions. Experiments on diverse synthetic and real-world benchmarks demonstrate better conditional calibration and substantially shorter intervals than existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。