arXiv:2510.23534econ.EMcs.LG2025-10被引 7

用Bregman散度统一框架,直接消除机器学习估计偏差。

Direct Debiased Machine Learning via Bregman Divergence Minimization

  • 通过Neyman正交得分最小化,端到端优化干扰参数与核函数。
  • 在因果推断中实现无偏估计,提升回归与密度比估计性能。
  • 适用于需要去偏的因果分析、结构模型,尤其适合高维数据。

我们提出一种直接去偏机器学习框架,结合Neyman目标估计与广义Riesz回归。该框架统一了Riesz回归、协变量平衡、目标最大似然估计(TMLE)和密度比估计。在涉及因果效应或结构模型的问题中,目标参数依赖于回归函数。直接使用机器学习估计的回归函数代入识别方程会产生第一阶段偏差。为降低偏差,去偏机器学习采用Neyman正交估计方程。传统方法需估计Riesz表示元与回归函数。为此,我们构建一个端到端算法:将干扰参数(回归函数与Riesz表示元)的估计建模为已知与未知干扰参数下计算的Neyman正交得分之间的差异最小化,称为Neyman目标估计。该方法包含Riesz表示元估计,使用Bregman散度衡量差异。该散度涵盖多种损失函数,平方损失对应Riesz回归,KL散度对应熵平衡。我们称此为广义Riesz回归。此外,该框架也包含TMLE作为回归函数估计的特例。对于特定模型对与表示元估计方法组合,可自动获得协变量平衡性质,无需显式求解平衡目标。

原文摘要 · Abstract (English)

We develop a direct debiased machine learning framework comprising Neyman targeted estimation and generalized Riesz regression. Our framework unifies Riesz regression for automatic debiased machine learning, covariate balancing, targeted maximum likelihood estimation (TMLE), and density-ratio estimation. In many problems involving causal effects or structural models, the parameters of interest depend on regression functions. Plugging regression functions estimated by machine learning methods into the identifying equations can yield poor performance because of first-stage bias. To reduce such bias, debiased machine learning employs Neyman orthogonal estimating equations. Debiased machine learning typically requires estimation of the Riesz representer and the regression function. For this problem, we develop a direct debiased machine learning framework with an end-to-end algorithm. We formulate estimation of the nuisance parameters, the regression function and the Riesz representer, as minimizing the discrepancy between Neyman orthogonal scores computed with known and unknown nuisance parameters, which we refer to as Neyman targeted estimation. Neyman targeted estimation includes Riesz representer estimation, and we measure discrepancies using the Bregman divergence. The Bregman divergence encompasses various loss functions as special cases, where the squared loss yields Riesz regression and the Kullback-Leibler divergence yields entropy balancing. We refer to this Riesz representer estimation as generalized Riesz regression. Neyman targeted estimation also yields TMLE as a special case for regression function estimation. Furthermore, for specific pairs of models and Riesz representer estimation methods, we can automatically obtain the covariate balancing property without explicitly solving the covariate balancing objective.

去偏学习因果推断机器学习统计估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。