arXiv:2603.05575stat.MLcs.LG2026-03被引 1

利用大量无标签数据和预训练模型,提升稀缺标注数据下条件推断的精度与效率。

Prediction-Powered Conditional Inference

  • 通过核方法自适应加权,将条件均值转化为加权无条件矩进行估计。
  • 引入预测修正项,在预测准确时显著降低方差,且不依赖预测精度保证有效性。
  • 适用于标注数据少、无标签数据多的场景,尤其适合已有黑箱模型的研究者。

在标注数据稀缺、无标签协变量丰富且已有黑箱机器学习预测器的场景下,研究预测驱动的条件推断。目标是对固定目标点的条件函数(如条件均值)进行统计推断,无需对条件关系施加参数化假设。方法结合局部化与预测驱动的方差缩减:首先提出基于RKHS的自适应加权局部化方法,将目标条件矩转化为加权无条件矩;其次通过修正分解形式融合机器学习预测结果,得到预测驱动的估计量与置信区间,在预测有效时可大幅降方差,且无论预测精度如何均保持推断有效性。建立了非渐近误差界,在无标签数据充足时达到极小极大最优收敛速率,证明了点态渐近正态性及方差估计一致性,并给出显式方差分解,揭示预测与无标签数据如何共同提升统计效率。模拟与真实数据实验表明,该方法具有有效的条件覆盖概率,且置信区间显著比现有方法更窄。

原文摘要 · Abstract (English)

We study prediction-powered conditional inference in the setting where labeled data are scarce, unlabeled covariates are abundant, and a black-box machine-learning predictor is available. The goal is to perform statistical inference on conditional functionals evaluated at a fixed target point, such as conditional means, without imposing a parametric model for the conditional relationship. Our approach combines localization with prediction-based variance reduction. First, we introduce an RKHS localization method that learns a data-adaptive weight from covariates and reformulates the target conditional moment at the target point as a weighted unconditional moment. Second, we incorporate machine-learning predictions through a correction-based decomposition of this localized moment, yielding a prediction-powered estimator and confidence interval that reduce variance when the predictor is informative while preserving validity regardless of predictor accuracy. We establish nonasymptotic error bounds and, in the abundant-unlabeled regime, minimax-optimal convergence rates for the resulting estimator, prove pointwise asymptotic normality with consistent variance estimation, and provide an explicit variance decomposition that characterizes how machine-learning predictions and unlabeled covariates improve statistical efficiency. Numerical experiments on simulated and real datasets demonstrate valid conditional coverage and substantially sharper confidence intervals than alternative methods.

条件推断预测驱动无标签数据统计效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。