arXiv:2506.05583cs.LGcs.AI2025-06被引 3

在未知群体分布变化下,仍能保证预测置信度的可靠性。

Conformal Prediction Adaptive to Unknown Subpopulation Shifts

  • 无需已知数据群体标签,自动适应未知分布偏移
  • 在视觉与语言任务中保持覆盖概率,标准方法失效时仍有效
  • 适用于高维场景,适合真实机器学习应用

置信预测广泛用于为黑箱机器学习模型提供不确定性量化,并在可交换数据下提供形式化覆盖保证。然而,当测试环境中的子群体混合比例与校准数据不同时,这些保证会失效。本文研究未知子群体偏移情形,即未提供子群体标签,需从数据中推断。提出新方法,可证明性地使置信预测适应此类偏移,无需显式知晓子群体结构。相比已有方法假设完美标签,本框架明确放宽该要求,并刻画了覆盖保证仍可行的条件。算法可扩展至高维设置,在真实任务中仍具实用性。在视觉(使用视觉变换器)和语言(使用大语言模型)基准上的大量实验表明,所提方法在标准置信预测失效的场景中,仍能可靠维持覆盖率并有效控制风险。

原文摘要 · Abstract (English)

Conformal prediction is widely used to equip black-box machine learning models with uncertainty quantification, offering formal coverage guarantees under exchangeable data. However, these guarantees fail when faced with subpopulation shifts, where the test environment contains a different mix of subpopulations than the calibration data. In this work, we focus on unknown subpopulation shifts where we are not given group-information i.e. the subpopulation labels of datapoints have to be inferred. We propose new methods that provably adapt conformal prediction to such shifts, ensuring valid coverage without explicit knowledge of subpopulation structure. While existing methods in similar setups assume perfect subpopulation labels, our framework explicitly relaxes this requirement and characterizes conditions where formal coverage guarantees remain feasible. Further, our algorithms scale to high-dimensional settings and remain practical in realistic machine learning tasks. Extensive experiments on vision (with vision transformers) and language (with large language models) benchmarks demonstrate that our methods reliably maintain coverage and effectively control risks in scenarios where standard conformal prediction fails.

置信预测分布偏移不确定性量化自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。