PPI++看似有免费午餐,但小样本下可能更差,关键看伪标签与真标签相关性。
No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered Inference
- 通过精确小样本分析,揭示PPI++的改进依赖于伪标签与真标签的相关性。
- 当伪标签相关性低于1/√(n−2)时,使用真标签反而更优,实证验证成立。
- 为实践者提供可透明评估PPI++收益的理论工具,适合数据标注不完美的场景。
预测驱动推断(PPI)是一种结合高质量真标签与可能噪声的伪标签进行统计估计的流行策略。以往研究显示,PPI++这一自适应形式在渐近意义下具有“免费午餐”特性:其渐近方差始终不大于仅使用真标签的方差,且该结果对伪标签质量无要求。本文通过针对均值估计问题的精确有限样本分析,揭示了这一现象背后的真相。我们提出“无免费午餐”结论,明确指出了在哪些样本量和设置下,PPI++的估计误差会显著劣于仅使用真标签。具体而言,只有当伪标签与真标签的相关性超过依赖于样本数n的阈值时,PPI++才可能表现更优。以高斯数据为例,相关性需至少达到1/√(n−2)才能带来改善。更广泛地,本文给出了在样本分割条件下PPI++方差的精确非渐近表达式,旨在帮助从业者在特定应用中透明判断其优势。实验表明,理论发现可在真实数据集上复现。
原文摘要 · Abstract (English)
Prediction-Powered Inference (PPI) is a popular strategy for combining gold-standard and possibly noisy pseudo-labels to perform statistical estimation. Prior work has shown an asymptotic \enquote{free lunch} for PPI++, an adaptive form of PPI, showing that the \textit{asymptotic} variance of PPI++ is always less than or equal to the variance obtained from using gold-standard labels alone. Notably, this result holds \textit{regardless of the quality of the pseudo-labels}. In this work, we demystify this result by conducting an exact finite-sample analysis of the estimation error of PPI++ on the mean estimation problem. We give a \enquote{no free lunch} result, characterizing the settings (and sample sizes) where PPI++ has provably worse estimation error than using gold-standard labels alone. Specifically, PPI++ will outperform if and only if the correlation between pseudo- and gold-standard is above a certain level that depends on the number of labeled samples ($n$). In some cases our results simplify considerably: For Gaussian data, for instance, the correlation must be at least $1/\sqrt{n - 2}$ in order to see improvement. More broadly, by providing exact non-asymptotic expressions for the variance of PPI++ under sample splitting, we aim to empower practitioners to transparently reason about the benefits of PPI++ in specific applications. In experiments, we illustrate that our theoretical findings hold on real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。