arXiv:2505.11785cs.LGcs.AI2025-05

加权融合多个预测集,提升不确定性量化精度

Improving Coverage in Combined Prediction Sets with Weighted p-values

  • 按贡献分配权重,动态融合多个预测集
  • 在实验中实现介于1-α与1-2α之间的自适应覆盖率
  • 支持数据依赖权重,适用于专家混合模型等场景

置信预测通过为点预测添加有效预测集来量化机器学习模型的不确定性。在涉及多个试验、模型或数据源的复杂场景中,可将多个预测集聚合以捕捉整体不确定性,通常能提高精度。然而,将多个具有单个1-α覆盖概率的预测集合并,不可避免地削弱整体保证,通常导致最坏情况下的1-2α覆盖率。本文提出一种加权预测集聚合框架,根据各预测集的贡献分配权重。该框架灵活控制聚合方式,实现介于组合模型的1-2α与单个模型的1-α之间的更紧覆盖边界,具体取决于权重分布。重要的是,该框架可推广至数据依赖权重:我们推导出一种保持有限样本有效性的加权聚合方法,即使权重依赖于数据也成立。这一扩展使框架广泛适用于权重可学习的场景,如专家混合(Mixture-of-Experts, MoE)。实验表明,在MoE设置下,该方法实现了自适应覆盖率。

原文摘要 · Abstract (English)

Conformal prediction quantifies the uncertainty of machine learning models by augmenting point predictions with valid prediction sets. For complex scenarios involving multiple trials, models, or data sources, conformal prediction sets can be aggregated to create a prediction set that captures the overall uncertainty, often improving precision. However, aggregating multiple prediction sets with individual $1-α$ coverage inevitably weakens the overall guarantee, typically resulting in $1-2α$ worst-case coverage. In this work, we propose a framework for the weighted aggregation of prediction sets, where weights are assigned to each prediction set based on their contribution. Our framework offers flexible control over how the sets are aggregated, achieving tighter coverage bounds that interpolate between the $1-2α$ guarantee of the combined models and the $1-α$ guarantee of an individual model depending on the distribution of weights. Importantly, our framework generalizes to data-dependent weights, as we derive a procedure for weighted aggregation that maintains finite-sample validity even when the weights depend on the data. This extension makes our framework broadly applicable to settings where weights are learned, such as mixture-of-experts (MoE), and we demonstrate through experiments in the MoE setting that our methods achieve adaptive coverage.

置信预测不确定性量化加权融合MoE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。