arXiv:2605.11638stat.MLcs.LG2026-05

通过主动推理优化标签选择,提升统计估计效率。

Learning U-Statistics with Active Inference

论文配图:Learning U-Statistics with Active Inference
图 1 · 摘自论文原文
  • 基于加权U统计量设计主动采样策略,融合模型预测与采样规则。
  • 在固定标注预算下,显著降低估计方差,保持目标覆盖率。
  • 适用于标注成本高、需高效统计推断的场景。

U-统计量在统计推断中具有核心作用。但在许多现代应用中,获取用于U-统计量的标签代价高昂。受主动推理最新进展启发,本文提出一种面向U-统计量的主动推理框架,通过选择性查询信息量大的标签,在固定标注预算下提升估计效率,同时保证有效的统计推断。方法基于增强的逆概率加权U统计量,旨在融合采样规则与机器学习预测。我们刻画了最小化其方差的最优采样规则,并设计了可实用的采样策略。进一步将该框架扩展至基于U-统计量的经验风险最小化。在真实数据集上的实验表明,相比基线方法,估计效率有显著提升,且保持目标覆盖率。

原文摘要 · Abstract (English)

$U$-statistics play a central role in statistical inference. In many modern applications, however, acquiring the labels required for $U$-statistics is costly. Motivated by recent advances in active inference, we develop an active inference framework for $U$-statistics that selectively queries informative labels to improve estimation efficiency under a fixed labeling budget, while preserving valid statistical inference. Our approach is built on the augmented inverse probability weighting $U$-statistic, which is designed to incorporate the sampling rule and machine learning predictions. We characterize the optimal sampling rule that minimizes its variance and design practical sampling strategies. We further extend the framework to $U$-statistic-based empirical risk minimization. Experiments on real datasets demonstrate substantial gains in estimation efficiency over baseline methods, while maintaining target coverage.

主动推理统计推断U统计量标注效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。