arXiv:2411.03195stat.MLcs.LG2024-11

在线选择数据源,动态分配预算以更准估计目标参数。

Online Data Collection for Efficient Semiparametric Inference

  • 基于在线矩选择框架,实时决定从哪个数据源采样。
  • 两种策略在真实数据上均实现零后悔,优于固定采样方案。
  • 适合需要边收集边决策的因果推断等场景。

尽管已有大量研究关注统计数据融合,但通常假设数据集是预先给定的。实际上,估计过程涉及复杂的数据收集决策,如确定可用数据源、其成本以及每个来源应采集的样本量。这一过程往往是顺序进行的,因为当前收集的数据可改善后续的决策。在本研究中,给定多个数据源和预算约束,智能体需顺序决定查询哪个数据源以高效估计目标参数。我们采用在线矩选择(Online Moment Selection)这一半参数框架,该框架适用于由一组矩条件定义的任意参数。有趣的是,最优预算分配依赖于未知的真实参数值。我们提出两种在线数据收集策略:探索后承诺(Explore-then-Commit)与探索后贪心(Explore-then-Greedy),它们利用当前参数估计值来最优地分配剩余预算。我们证明,这两种策略相对于一个已知真值的基准策略,实现了渐近均方误差意义上的零后悔。我们在合成数据和真实世界因果效应估计任务上进行了实证验证,结果表明在线数据收集策略显著优于固定采样方法。

原文摘要 · Abstract (English)

While many works have studied statistical data fusion, they typically assume that the various datasets are given in advance. However, in practice, estimation requires difficult data collection decisions like determining the available data sources, their costs, and how many samples to collect from each source. Moreover, this process is often sequential because the data collected at a given time can improve collection decisions in the future. In our setup, given access to multiple data sources and budget constraints, the agent must sequentially decide which data source to query to efficiently estimate a target parameter. We formalize this task using Online Moment Selection, a semiparametric framework that applies to any parameter identified by a set of moment conditions. Interestingly, the optimal budget allocation depends on the (unknown) true parameters. We present two online data collection policies, Explore-then-Commit and Explore-then-Greedy, that use the parameter estimates at a given time to optimally allocate the remaining budget in the future steps. We prove that both policies achieve zero regret (assessed by asymptotic MSE) relative to an oracle policy. We empirically validate our methods on both synthetic and real-world causal effect estimation tasks, demonstrating that the online data collection policies outperform their fixed counterparts.

在线学习半参数估计因果推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。