arXiv:2411.12036stat.MLcs.LG2024-11被引 5

用预测引导实验采样,提升效率与精度。

Prediction-Guided Active Experiments

  • 基于预测结果动态选择实验样本,优化采样策略。
  • 理论证明可达到半参数效率下界,渐近最优。
  • 适用于需要高效实验设计的科研与政策评估场景。

本文提出一种新的主动实验框架——预测引导主动实验(Prediction-Guided Active Experiment, PGAE),利用现有机器学习模型的预测结果指导实验采样与观测。在每个时间步,根据指定采样分布选取实验单元,并依据实验概率观察实际结果;否则仅能获得预测值。首先分析非自适应情形,假设已知预测变量与真实结果的联合分布,通过最小化正则估计量的半参数效率界,推导出最优实验策略,并提出一种可达到该效率界的估计器,实现渐近最优。随后拓展至自适应情形,即预测模型随新采样数据持续更新,在一定正则性条件下,该自适应估计器仍保持效率并达到相同半参数界。最后通过模拟实验及基于美国人口普查局数据的半合成实验验证了PGAE的有效性,结果表明其性能优于现有方法。

原文摘要 · Abstract (English)

In this work, we introduce a new framework for active experimentation, the Prediction-Guided Active Experiment (PGAE), which leverages predictions from an existing machine learning model to guide sampling and experimentation. Specifically, at each time step, an experimental unit is sampled according to a designated sampling distribution, and the actual outcome is observed based on an experimental probability. Otherwise, only a prediction for the outcome is available. We begin by analyzing the non-adaptive case, where full information on the joint distribution of the predictor and the actual outcome is assumed. For this scenario, we derive an optimal experimentation strategy by minimizing the semi-parametric efficiency bound for the class of regular estimators. We then introduce an estimator that meets this efficiency bound, achieving asymptotic optimality. Next, we move to the adaptive case, where the predictor is continuously updated with newly sampled data. We show that the adaptive version of the estimator remains efficient and attains the same semi-parametric bound under certain regularity assumptions. Finally, we validate PGAE's performance through simulations and a semi-synthetic experiment using data from the US Census Bureau. The results underscore the PGAE framework's effectiveness and superiority compared to other existing methods.

主动实验预测引导效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。