在分布偏移下,用源域全信息+目标域特征,精准估计最优策略。
Optimal Policy Adaptation under Covariate Shift
- 基于因果视角建模分布偏移,推导奖励可识别性假设。
- 提出双重稳健且半参数高效的奖励估计器,显著提升准确率。
- 理论分析偏差与泛化误差,适合需可靠策略迁移的场景。
预测模型的迁移学习已得到广泛研究,但相应的策略学习方法却很少被讨论。本文提出一种在目标域中学习最优策略的系统性方法,利用两个数据集:一个来自源域的完整信息数据集,另一个仅包含目标域的协变量。在协变量偏移设定下,从因果角度建模问题,并给出由给定策略诱导的奖励的可识别性假设。进一步推导出奖励的高效影响函数和半参数效率界。基于此,构建了一个双重稳健且半参数高效的奖励估计器,并通过优化估计奖励来学习最优策略。此外,对所学策略的偏差和泛化误差边界进行了理论分析。大量实验表明,该方法不仅更准确地估计奖励,而且生成的策略能紧密逼近理论最优策略。
原文摘要 · Abstract (English)
Transfer learning of prediction models has been extensively studied, while the corresponding policy learning approaches are rarely discussed. In this paper, we propose principled approaches for learning the optimal policy in the target domain by leveraging two datasets: one with full information from the source domain and the other from the target domain with only covariates. First, under the setting of covariate shift, we formulate the problem from a perspective of causality and present the identifiability assumptions for the reward induced by a given policy. Then, we derive the efficient influence function and the semiparametric efficiency bound for the reward. Based on this, we construct a doubly robust and semiparametric efficient estimator for the reward and then learn the optimal policy by optimizing the estimated reward. Moreover, we theoretically analyze the bias and the generalization error bound for the learned policy. Extensive experiments demonstrate that the approach not only estimates the reward more accurately but also yields a policy that closely approximates the theoretically optimal policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。