揭示政策学习中两种主流方法的数学等价性,为优化提供新思路。
Bridging the Gap between Empirical Welfare Maximization and Conditional Average Treatment Effect Estimation in Policy Learning
- 发现经验福利最大化与条件平均处理效应估计本质同源
- 理论证明二者在重参数化下可等价转化为最小二乘问题
- 提出正则化策略,提升实际优化效率,适合政策制定者参考
政策学习的目标是训练一个策略函数,在给定协变量时推荐最优处理以最大化总体福利。目前主要有两种方法:经验福利最大化(EWM)和插件法。EWM将问题视为分类任务,先估计群体福利函数,再通过最大化估计值训练策略;插件法则基于回归,先估计条件平均处理效应(CATE),再选择预测结果最高的处理。本文揭示了这两种方法的本质一致性,证明在策略类的重参数化下,EWM等价于最小二乘优化。由此,两者在常见条件下具有相同的理论保障,且可互换使用。基于此等价性,本文提出一种正则化方法,将问题转化为更易优化的光滑代理目标。然而,对于许多自然策略类,精确EWM的组合复杂性仍存在,因此该转化主要作为优化辅助工具,而非对NP难问题的通用解法。
原文摘要 · Abstract (English)
The goal of policy learning is to train a policy function that recommends a treatment given covariates to maximize population welfare. There are two major approaches in policy learning: the empirical welfare maximization (EWM) approach and the plug-in approach. The EWM approach is analogous to a classification problem, where one first builds an estimator of the population welfare, which is a functional of policy functions, and then trains a policy by maximizing the estimated welfare. In contrast, the plug-in approach is based on regression, where one first estimates the conditional average treatment effect (CATE) and then recommends the treatment with the highest estimated outcome. This study bridges the gap between the two approaches by showing that both are based on essentially the same optimization problem. In particular, we prove an exact equivalence between EWM and least squares over a reparameterization of the policy class. As a consequence, the two approaches are interchangeable in several respects and share the same theoretical guarantees under common conditions. Leveraging this equivalence, we propose a regularization method for policy learning. The reduction to least squares yields a smooth surrogate that is typically easier to optimize in practice. At the same time, for many natural policy classes the inherent combinatorial hardness of exact EWM generally remains, so the reduction should be viewed as an optimization aid rather than a universal bypass of NP-hardness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。