通过自适应查询策略,提升均值估计的精度与可靠性。
Revisiting Active Sequential Prediction-Powered Mean Estimation
- 结合不确定性与固定概率,动态调整标签查询策略。
- 理论证明置信区间受数据依赖约束,性能更优。
- 适合关注主动学习与统计推断交叉领域的研究者。
本文重新审视了主动序列预测驱动的均值估计问题:在每一轮中,需根据样本协变量决定是否查询真实标签的概率。若不查询,则使用机器学习模型的预测结果替代。已有方法通过将基于不确定性的建议与固定概率相结合来设定查询概率,其中固定概率体现对查询频率的软约束。我们探索了混合参数的不同取值,发现当固定概率权重接近1时,置信区间最窄,不确定性成分影响最小。受此启发,我们建立了该估计器的非渐近分析,推导出数据依赖的置信区间上界。进一步分析表明,若采用无遗憾学习方法确定查询概率并控制该上界,则查询概率会收敛至一个全局最大值(即忽略当前协变量时的最优查询概率)。模拟实验验证了理论结果。
原文摘要 · Abstract (English)
In this work, we revisit the problem of active sequential prediction-powered mean estimation, where at each round one must decide the query probability of the ground-truth label upon observing the covariates of a sample. Furthermore, if the label is not queried, the prediction from a machine learning model is used instead. Prior work proposed an elegant scheme that determines the query probability by combining an uncertainty-based suggestion with a constant probability that encodes a soft constraint on the query probability. We explored different values of the mixing parameter and observed an intriguing empirical pattern: the smallest confidence width tends to occur when the weight on the constant probability is close to one, thereby reducing the influence of the uncertainty-based component. Motivated by this observation, we develop a non-asymptotic analysis of the estimator and establish a data-dependent bound on its confidence interval. Our analysis further suggests that when a no-regret learning approach is used to determine the query probability and control this bound, the query probability converges to the constraint of the max value of the query probability when it is chosen obliviously to the current covariates. We also conduct simulations that corroborate these theoretical findings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。