arXiv:2605.12341stat.MLcs.LG2026-05

新方法让预测集更小更稳,不用分数据也能保证准确率。

Multi-Variable Conformal Prediction: Optimizing Prediction Sets without Data Splitting

论文配图:Multi-Variable Conformal Prediction: Optimizing Prediction Sets without Data Splitting
图 1 · 摘自论文原文
  • 用多个评分变量同时优化预测集形状,不再依赖数据分割。
  • 在椭球和多模态场景中,覆盖率达标且集合尺寸更小。
  • 适合追求高可靠性与低方差的机器学习应用者。

传统置信推断通过数据分割实现有限样本覆盖保证,但其校准阶段受限于标量评分函数与单一阈值变量,导致预测集形状固定。本文提出多变量置信推断(MCP),将向量评分函数与多个同步校准变量引入框架,基于情景理论将预测集设计与校准统一为单个优化问题,无需数据分割即可保持覆盖保证。提出两种高效变体:RemMCP基于约束优化与约束移除,可推广分裂置信推断;RelMCP采用迭代优化与约束松弛,支持非凸评分函数但可能更保守。在椭球与多模态预测集上的实验表明,两种方法均稳定达到目标覆盖度,预测集尺寸小于或相当于基线方法,且校准运行间方差显著降低——源于全部数据用于形状优化与校准的联合过程。

原文摘要 · Abstract (English)

Conformal prediction constructs prediction sets with finite-sample coverage guarantees, but its calibration stage is structurally constrained to a scalar score function and a single threshold variable - forcing shapes of prediction sets to be fixed before calibration, typically through data splitting. We introduce multi-variable conformal prediction (MCP), a framework that extends conformal prediction to vector-valued score functions with multiple simultaneous calibration variables. Building on scenario theory as a principled framework for certifying data-driven decisions, MCP unifies prediction set design and calibration into a single optimization problem, eliminating data splitting without sacrificing coverage guarantees. We propose two computationally efficient variants: RemMCP, grounded in constrained optimization with constraint removal, which admits a clean generalization of split conformal prediction; and RelMCP, based on iterative optimization with constraint relaxation, which supports non-convex score functions at the cost of possibly greater conservatism. Through numerical experiments on ellipsoidal and multi-modal prediction sets, we demonstrate that RemMCP and RelMCP consistently meet the target coverage with prediction set sizes smaller than or comparable to those of baselines with data split, while considerably reducing variance across calibration runs - a direct consequence of using all available data for shape optimization and calibration simultaneously.

置信推断预测集无数据分割优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。