针对连续干预政策评估,提出抗分布偏移的鲁棒学习方法
Distributionally Robust Policy Evaluation and Learning for Continuous Treatment with Observational Data
- 用核函数改进IPW法,缓解连续处理下的观测遗漏问题
- 在存在分布偏移时仍保证估计器收敛性,理论严谨
- 适合医疗、经济等存在连续干预与数据偏移的真实场景
利用离线观测数据进行政策评估与学习,使决策者能够建立特征与干预之间的策略。现有研究多集中于离散干预空间,或假设策略学习与部署环境分布无差异,这限制了在存在分布偏移的连续干预真实场景中的应用。本文针对连续处理设置,提出一种分布鲁棒策略评估与学习方法。通过将核函数引入扩展的逆概率加权(IPW)估计器,缓解标准IPW方法在连续处理中对部分观测的排除问题。进一步提供有限样本分析,证明所提分布鲁棒估计器的收敛性。全面实验验证了该方法在存在分布偏移时的有效性。
原文摘要 · Abstract (English)
Using offline observational data for policy evaluation and learning allows decision-makers to evaluate and learn a policy that connects characteristics and interventions. Most existing literature has focused on either discrete treatment spaces or assumed no difference in the distributions between the policy-learning and policy-deployed environments. These restrict applications in many real-world scenarios where distribution shifts are present with continuous treatment. To overcome these challenges, this paper focuses on developing a distributionally robust policy under a continuous treatment setting. The proposed distributionally robust estimators are established using the Inverse Probability Weighting (IPW) method extended from the discrete one for policy evaluation and learning under continuous treatments. Specifically, we introduce a kernel function into the proposed IPW estimator to mitigate the exclusion of observations that can occur in the standard IPW method to continuous treatments. We then provide finite-sample analysis that guarantees the convergence of the proposed distributionally robust policy evaluation and learning estimators. The comprehensive experiments further verify the effectiveness of our approach when distribution shifts are present.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。