高维数据下实现精准预测集,无需直接估计似然比。
Conformal Inference under High-Dimensional Covariate Shifts via Likelihood-Ratio Regularization
- 用带正则化的分位数回归间接构造阈值函数
- 在高维图像和文本任务中覆盖率达目标水平且误差可控
- 适合处理分布偏移下的可靠预测,尤其适合复杂数据
在协变量偏移下研究置信推断问题。给定来自源域的标注数据和目标域的未标注数据,目标是构建在目标域中具有有效边际覆盖率的预测集。现有方法多需估计未知的似然比函数,对高维数据(如图像)而言代价高昂。为此,本文提出似然比正则化分位数回归(LR-QR)算法,将pinball损失与新型正则化结合,无需直接估计未知似然比即可构造阈值函数。理论证明该方法在目标域中的覆盖率接近所需水平,误差项可控制。证明基于学习理论中的稳定性边界分析。实验表明,LR-QR在多个高维任务中表现优于现有方法:包括针对Communities and Crime数据集的回归任务、来自WILDS仓库的图像分类任务,以及MMLU基准上的大语言模型问答任务。
原文摘要 · Abstract (English)
We consider the problem of conformal prediction under covariate shift. Given labeled data from a source domain and unlabeled data from a covariate shifted target domain, we seek to construct prediction sets with valid marginal coverage in the target domain. Most existing methods require estimating the unknown likelihood ratio function, which can be prohibitive for high-dimensional data such as images. To address this challenge, we introduce the likelihood ratio regularized quantile regression (LR-QR) algorithm, which combines the pinball loss with a novel choice of regularization in order to construct a threshold function without directly estimating the unknown likelihood ratio. We show that the LR-QR method has coverage at the desired level in the target domain, up to a small error term that we can control. Our proofs draw on a novel analysis of coverage via stability bounds from learning theory. Our experiments demonstrate that the LR-QR algorithm outperforms existing methods on high-dimensional prediction tasks, including a regression task for the Communities and Crime dataset, an image classification task from the WILDS repository, and an LLM question-answering task on the MMLU benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。