解决隐含混杂因素下的个性化策略学习偏差问题
Efficient and Sharp Off-Policy Learning under Unobserved Confounding
- 基于因果敏感性分析,构建高效且稳定的策略价值边界估计器
- 在合成与真实数据上优于传统方法,显著降低策略偏差
- 适合医疗决策、公共政策等存在隐性干扰的场景
我们提出一种新方法,用于存在未观测混杂因素时的个性化离线策略学习。标准策略学习依赖于无混杂假设(即无未观测因素同时影响处理分配和结果),但该假设常被违背,导致估计偏差并产生有害策略。为此,我们引入因果敏感性分析,推导出在未观测混杂下价值函数的紧边界半参数高效估计量。该估计量具有三重优势:(1) 避免基于逆倾向得分加权结果的不稳定性最小最大优化;(2) 具备半参数效率;(3) 证明其可生成最优抗混杂策略。进一步将理论拓展至有基准策略(如标准治疗)时的策略改进任务。实验表明,我们的方法在合成与真实数据上均优于简单插补法及现有基线。本方法对医疗、公共政策等存在未观测混杂的决策场景具有高度适用性。
原文摘要 · Abstract (English)
We develop a novel method for personalized off-policy learning in scenarios with unobserved confounding. Thereby, we address a key limitation of standard policy learning: standard policy learning assumes unconfoundedness, meaning that no unobserved factors influence both treatment assignment and outcomes. However, this assumption is often violated, because of which standard policy learning produces biased estimates and thus leads to policies that can be harmful. To address this limitation, we employ causal sensitivity analysis and derive a semi-parametrically efficient estimator for a sharp bound on the value function under unobserved confounding. Our estimator has three advantages: (1) Unlike existing works, our estimator avoids unstable minimax optimization based on inverse propensity weighted outcomes. (2) Our estimator is semi-parametrically efficient. (3) We prove that our estimator leads to the optimal confounding-robust policy. Finally, we extend our theory to the related task of policy improvement under unobserved confounding, i.e., when a baseline policy such as the standard of care is available. We show in experiments with synthetic and real-world data that our method outperforms simple plug-in approaches and existing baselines. Our method is highly relevant for decision-making where unobserved confounding can be problematic, such as in healthcare and public policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。