arXiv:2512.17341stat.MLcs.LG2025-12被引 1

提出线性泛函估计的最优性理论,揭示双重稳健估计在因果推断中的根本优势。

Sharp Structure-Agnostic Lower Bounds for General Linear Functional Estimation

  • 基于黑箱估计器框架,不假设干扰函数结构,仅依赖非参数回归/分类器
  • 证明双重稳健估计在平均处理效应中达到统计最优,且一阶去偏在两类场景下均最优
  • 适用于广义回归和协变量偏移等复杂场景,为因果推断提供理论依据

我们建立了针对未知干扰分量的线性泛函估计问题的一般统计最优性理论。该框架涵盖许多因果和预测参数,具有广泛的应用价值。采用Balakrishnan等人提出的结构无关框架,不假设干扰函数的结构特性,仅要求存在能实现特定统计估计速率的黑箱估计器。该框架特别适合仅使用非参数回归与分类器作为黑箱子过程的估计策略。我们首先证明了在平均处理效应(ATE)估计中,被广泛使用的双重稳健估计器具有统计最优性。随后,我们刻画了该一般形式下的极小极大最优率。值得注意的是,我们区分了双重稳健性可实现与不可实现的两种情形,并指出一阶去偏在不同情形下产生不同的误差率。结果表明,一阶去偏在两类情形下均达到最优。我们通过实例化理论,推导出最优误差率,既恢复了已有结果,也扩展至多种感兴趣的情形,包括干扰由广义回归定义、训练与测试分布存在协变量偏移等情况。

原文摘要 · Abstract (English)

We establish a general statistical optimality theory for estimation problems where the target parameter is a linear functional of an unknown nuisance component that must be estimated from data. This formulation covers many causal and predictive parameters and has applications to numerous disciplines. We adopt the structure-agnostic framework introduced by \citet{balakrishnan2023fundamental}, which poses no structural properties on the nuisance functions other than access to black-box estimators that achieve some statistical estimation rate. This framework is particularly appealing when one is only willing to consider estimation strategies that use non-parametric regression and classification oracles as black-box sub-processes. Within this framework, we first prove the statistical optimality of the celebrated and widely used doubly robust estimators for the Average Treatment Effect (ATE), the most central parameter in causal inference. We then characterize the minimax optimal rate under the general formulation. Notably, we differentiate between two regimes in which double robustness can and cannot be achieved and in which first-order debiasing yields different error rates. Our result implies that first-order debiasing is simultaneously optimal in both regimes. We instantiate our theory by deriving optimal error rates that recover existing results and extend to various settings of interest, including the case when the nuisance is defined by generalized regressions and when covariate shift exists for training and test distribution.

因果推断统计最优性双重稳健线性泛函

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。