用置信预测扩展黑箱模型的缺失数据推断,同时保障隐私和稳定性。
Extending Prediction-Powered Inference through Conformal Prediction
- 通过校准的置信集预测器进行数据填补,实现推断有效性
- 在均值、Z-估计等任务中保持推断准确性和额外鲁棒性
- 首次支持离线运行的e值推断,适用于隐私与时间序列数据
预测驱动推断是一种用于安全使用黑箱机器学习模型填补缺失数据的新方法,可强化统计参数的推断能力。然而,许多应用场景还需满足隐私保护、鲁棒性或分布持续漂移下的有效性等强约束,而现有方法需逐案设计且过程复杂。本文通过将预测驱动推断与置信预测相结合:利用校准的置信集预测器进行数据填补,自然地在保证推断有效性的同时获得额外性质。我们将其应用于均值、Z-和M-估计,以及e值与基于e值的推断过程。尤其在e值场景下,本方法是首个可在离线条件下运行的通用预测驱动方案。实验表明,在隐私数据和时间序列数据上的应用均具挑战性,但在本框架下变得自然可行。
原文摘要 · Abstract (English)
Prediction-powered inference is a recent methodology for the safe use of black-box ML models to impute missing data, strengthening inference of statistical parameters. However, many applications require strong properties besides valid inference, such as privacy, robustness or validity under continuous distribution shifts; deriving prediction-powered methods with such guarantees is generally an arduous process, and has to be done case by case. In this paper, we resolve this issue by connecting prediction-powered inference with conformal prediction: by performing imputation through a calibrated conformal set-predictor, we attain validity while achieving additional guarantees in a natural manner. We instantiate our procedure for the inference of means, Z- and M-estimation, as well as e-values and e-value-based procedures. Furthermore, in the case of e-values, ours is the first general prediction-powered procedure that operates off-line. We demonstrate these advantages by applying our method on private and time-series data. Both tasks are nontrivial within the standard prediction-powered framework but become natural under our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。