arXiv:2601.21455stat.MLcs.LG2026-01

揭示置信区间越短未必越好,提出新指标检测误导性优化

Questioning the Coverage-Length Metric in Conformal Prediction: When Shorter Intervals Are Not Better

  • 用概率性空区间技巧伪装缩短区间长度,覆盖度仍达标
  • 同一输入多次运行结果不一致,暴露算法实际不稳定
  • 引入区间稳定性指标,可识别类似‘欺骗性’优化方法

置信预测(Conformal Prediction, CP)是分布无关不确定性量化的核心方法,传统上通过覆盖率和区间长度评估。本文质疑这两个标准指标的充分性。我们发现,通过一种反直觉的‘偏见技巧’(Prejudicial Trick, PT),可使区间长度看似显著缩短,而覆盖率依然有效。具体而言,对任意测试样本,PT以概率返回空区间或基于调整置信水平构造的区间,从而保持边际覆盖率。尽管长度被误导性地降低,但该方法导致严重实践缺陷:相同输入在多次运行中产生完全不同预测区间。我们形式化推导了PT实现误导性改进的条件,并在多种回归与分类任务中提供充分实证支持。此外,我们提出新指标‘区间稳定性’,用于检测新方法是否隐含采用类似PT的技术。代码已开源。

原文摘要 · Abstract (English)

Conformal prediction(CP) has become a cornerstone of distribution-free uncertainty quantification, conventionally evaluated by its coverage and interval length. This work critically examines the sufficiency of these standard metrics. We demonstrate that the interval length might be deceptively improved through a counter-intuitive approach termed Prejudicial Trick(PT), while the coverage remains valid. Specifically, for any given test sample, PT probabilistically returns an interval, which is either null or constructed using an adjusted confidence level, thereby preserving marginal coverage. While PT potentially yields a deceptively lower interval length, it introduces practical vulnerabilities: the same input can yield completely different prediction intervals across repeated runs of the algorithm. We formally derive the conditions under which PT achieves these misleading improvements and provide extensive empirical evidence across various regression and classification tasks. Furthermore, we introduce a new metric interval stability which helps detect whether a new CP method implicitly improves the length based on such PT-like techniques. Code is available at https://github.com/benben-cd/PT-Conformal-Prediction.

置信预测不确定性量化算法可靠性评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。