将因果森林改进为端到端政策学习算法,直接优化治疗推荐效果。
Causal-Policy Forest for End-to-End Policy Learning
- 基于CATE误差最小化重构政策学习目标,改进因果森林结构。
- 无需分步估计干扰项,实现政策参数的端到端联合训练。
- 保持树模型高效性,适合大规模个体化决策场景。
本研究提出一种用于因果推断中策略学习的端到端算法。观测数据包含协变量、处理分配和结果,仅可观测对应处理的结果。策略学习的目标是根据观测数据训练一个策略函数,为每个个体推荐最优处理以最大化策略价值。本文首次证明,在{-1, 1}受限回归模型下,最大化策略价值等价于最小化条件平均处理效应(CATE)的均方误差。基于此发现,我们对因果森林——一种广泛使用的端到端CATE估计方法——进行了改造,提出因果策略森林(Causal-Policy Forest)。该算法具备三大优势:一是对现有主流CATE估计方法进行简单修改,有助于弥合策略学习与CATE估计之间的实践鸿沟;二是避免传统方法中将干扰项估计作为独立任务,实现更彻底的端到端训练;三是继承决策树与随机森林的高效训练机制,有效避免计算不可行性。
原文摘要 · Abstract (English)
This study proposes an end-to-end algorithm for policy learning in causal inference. We observe data consisting of covariates, treatment assignments, and outcomes, where only the outcome corresponding to the assigned treatment is observed. The goal of policy learning is to train a policy from the observed data, where a policy is a function that recommends an optimal treatment for each individual, to maximize the policy value. In this study, we first show that maximizing the policy value is equivalent to minimizing the mean squared error for the conditional average treatment effect (CATE) under $\{-1, 1\}$ restricted regression models. Based on this finding, we modify the causal forest, an end-to-end CATE estimation algorithm, for policy learning. We refer to our algorithm as the causal-policy forest. Our algorithm has three advantages. First, it is a simple modification of an existing, widely used CATE estimation method, therefore, it helps bridge the gap between policy learning and CATE estimation in practice. Second, while existing studies typically estimate nuisance parameters for policy learning as a separate task, our algorithm trains the policy in a more end-to-end manner. Third, as in standard decision trees and random forests, we train the models efficiently, avoiding computational intractability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。