arXiv:2602.06297stat.MLcs.LG2026-02

提出可随时验证的预测方法,让机器学习在流式数据中也能可靠给出置信区间。

Time-uniform conformal and PAC prediction

  • 设计了无需预设样本量的序列化置信预测框架
  • 任意时刻的预测集覆盖率均满足设定水平
  • 适合动态决策场景,如金融风控、医疗监测

随着机器学习在高风险决策中的广泛应用,围绕黑箱模型的不确定性量化方法(如置信预测)受到广泛关注。在数据流式到达的序列场景中,传统置信方法需预先固定样本量,且无法处理不断更新的预测结果。为此,本文将置信预测与可能近似正确(PAC)预测框架扩展至序列设置,其中数据点数量无需事先确定。所提出的预测集具备即时有效性:分析师可在任意时刻选择查看结果,即使该选择依赖于数据,其期望覆盖率仍保持在指定水平。本文提供了理论保证,并在模拟与真实数据集上验证了方法的有效性与实用性。

原文摘要 · Abstract (English)

Given that machine learning algorithms are increasingly being deployed to aid in high stakes decision-making, uncertainty quantification methods that wrap around these black box models such as conformal prediction have received much attention in recent years. In sequential settings, where data are observed/generated in a streaming fashion, traditional conformal methods do not provide any guarantee without fixing the sample size. More importantly, traditional conformal methods cannot cope with sequentially updated predictions. As such, we develop an extension of the conformal prediction and related probably approximately correct (PAC) prediction frameworks to sequential settings where the number of data points is not fixed in advance. The resulting prediction sets are anytime-valid in that their expected coverage is at the required level at any time chosen by the analyst even if this choice depends on the data. We present theoretical guarantees for our proposed methods and demonstrate their validity and utility on simulated and real datasets.

置信预测序列分析不确定性量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。