用少量真实标签提升伪标签数据的统计推断精度,解决小样本下模型失效问题。
Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc Regression
- 将伪标签推断转化为后验回归,利用稳健回归降低方差
- 在仅少量真实标签时,新方法比传统PPI++更稳定可靠
- 适合标签稀缺场景下的模型评估与数据分析
可执行任意任务的机器学习系统生成的合成标签正被广泛用于统计推断,如数据分析或模型评估。预测驱动推断(PPI)框架通过结合大量伪标签数据与少量高质真实标签,实现低方差、无偏估计。然而,现有PPI研究多基于相对充足的标签,获取成本较高。本文发现当真实标签稀少时,PPI++反而可能劣于经典推断方法。我们通过将PPI++关联到普通最小二乘回归,揭示其在小样本下方差过高的本质原因,并基于此提出两种新方法:利用稳健回归器在少标签条件下进一步降低估计方差,显著提升推断效能。
原文摘要 · Abstract (English)
The availability of machine learning systems that can effectively perform arbitrary tasks has led to synthetic labels from these systems being used in applications of statistical inference, such as data analysis or model evaluation. The Prediction Powered Inference (PPI) framework provides a way of leveraging both a large pool of pseudo-labelled data and a small sample with real, high-quality labels to produce a low-variance, unbiased estimate of the quantity being evaluated for. Most work on PPI considers a relatively sizable set of labelled samples, which can be resource intensive to obtain. However, we find that when labelled data is scarce, the PPI++ method can perform even worse than classical inference. We analyze this phenomenon by relating PPI++ to ordinary least squares regression, which also experiences high variance with small sample sizes, and use this regression framework to better understand the efficacy of PPI. Motivated by this, we present two new PPI-based techniques that leverage robust regressors to produce even lower variance estimators in the few-label regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。