提出非渐近分析框架,揭示预测集效率受数据量与误差率的联合影响。
Non-Asymptotic Analysis of Efficiency in Conformalized Regression
- 基于非渐近理论,推导出预测集长度与理想区间的偏差上界。
- 发现效率随训练集大小n、校准集大小m和误差率α呈现分段收敛特性。
- 为实际应用中数据分配提供依据,适合关注可靠性与精度平衡的研究者。
置信预测可提供覆盖率保证。其信息量取决于效率,通常以预测集的期望长度衡量。以往研究常将误覆盖水平 α 视为固定常数。本文在数据分布满足弱假设下,针对通过 SGD 训练的分位数与中位数回归的置信化回归,建立了预测集长度偏离理想区间长度的非渐近界。该界为 𝒪(1/√n + 1/(α²n) + 1/√m + exp(−α²m)),捕捉了效率对训练集大小 n、校准集大小 m 以及误覆盖水平 α 的联合依赖关系。结果揭示了 α 不同取值区间下的收敛率相变现象,为控制预测集过长提供了数据分配指导。实验结果与理论一致。
原文摘要 · Abstract (English)
Conformal prediction provides prediction sets with coverage guarantees. The informativeness of conformal prediction depends on its efficiency, typically quantified by the expected size of the prediction set. Prior work on the efficiency of conformalized regression commonly treats the miscoverage level $α$ as a fixed constant. In this work, we establish non-asymptotic bounds on the deviation of the prediction set length from the oracle interval length for conformalized quantile and median regression trained via SGD, under mild assumptions on the data distribution. Our bounds of order $\mathcal{O}(1/\sqrt{n} + 1/(α^2 n) + 1/\sqrt{m} + \exp(-α^2 m))$ capture the joint dependence of efficiency on the proper training set size $n$, the calibration set size $m$, and the miscoverage level $α$. The results identify phase transitions in convergence rates across different regimes of $α$, offering guidance for allocating data to control excess prediction set length. Empirical results are consistent with our theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。