arXiv:2603.26415cs.LGcs.AI2026-03

解决数据分布偏移下的不确定性量化问题,提升模型可靠性。

KMM-CP: Practical Conformal Prediction under Covariate Shift via Selective Kernel Mean Matching

  • 用核均值匹配修正协变量偏移,稳定校准过程
  • 在分子属性预测中覆盖误差降低超50%
  • 适合高风险领域如医疗、科研中的可信推理

不确定性量化对科学发现和医疗等高风险领域至关重要。共形预测(CP)在可交换性假设下提供有限样本覆盖保证,但实际中常因分布偏移而失效。在协变量偏移下,恢复有效性需重要性加权,然而当训练与测试分布支持重叠有限时,密度比估计会不稳定。本文提出基于核均值匹配(KMM)的KMM-CP框架,通过在显式权重约束下最小化再生核希尔伯特空间(RKHS)矩差异,直接控制共形覆盖误差的偏差-方差成分,并在弱条件下建立渐近覆盖保证。进一步引入选择性扩展,识别可靠支持重叠区域,仅在此子集上进行共形校正,显著提升低重叠场景下的稳定性。在存在真实分布偏移的分子属性预测基准上,相比现有方法,KMM-CP覆盖差距减少超过50%。代码已开源:https://github.com/siddharthal/KMM-CP。

原文摘要 · Abstract (English)

Uncertainty quantification is essential for deploying machine learning models in high-stakes domains such as scientific discovery and healthcare. Conformal Prediction (CP) provides finite-sample coverage guarantees under exchangeability, an assumption often violated in practice due to distribution shift. Under covariate shift, restoring validity requires importance weighting, yet accurate density-ratio estimation becomes unstable when training and test distributions exhibit limited support overlap. We propose KMM-CP, a conformal prediction framework based on Kernel Mean Matching (KMM) for covariate-shift correction. We show that KMM directly controls the bias-variance components governing conformal coverage error by minimizing RKHS moment discrepancy under explicit weight constraints, and establish asymptotic coverage guarantees under mild conditions. We then introduce a selective extension that identifies regions of reliable support overlap and restricts conformal correction to this subset, further improving stability in low-overlap regimes. Experiments on molecular property prediction benchmarks with realistic distribution shifts show that KMM-CP reduces coverage gap by over 50% compared to existing approaches. The code is available at https://github.com/siddharthal/KMM-CP.

共形预测分布偏移不确定性量化分子性质

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。