用工具变量解决模仿学习中的隐藏干扰问题,提升策略估计准确性。
Confounded Causal Imitation Learning with Instrumental Variables
- 引入工具变量识别机制,破解多时步干扰导致的偏差
- 两阶段框架:先找有效工具变量,再优化策略
- 适用于存在隐藏变量的现实场景,如机器人控制
从示范中进行模仿学习通常受到未观测变量(即未测量混杂因子)对状态和动作的混淆影响。若忽略这些因素,会导致策略估计产生偏差。为打破这种混淆困境,本文结合工具变量(IV)的强大能力,提出一种受混杂影响的因果模仿学习(C2L)模型。该模型可处理跨多个时间步影响动作的混杂因子,而不仅限于即时时间依赖。我们构建了两阶段模仿学习框架,实现有效的工具变量识别与策略优化。第一阶段基于定义的伪变量构造检验准则,从而识别出满足充分必要可识别条件的有效工具变量。第二阶段利用已识别的工具变量,提出两种候选策略学习方法:一种基于模拟器,另一种为离线学习。大量实验验证了有效工具变量识别及策略学习的有效性。
原文摘要 · Abstract (English)
Imitation learning from demonstrations usually suffers from the confounding effects of unmeasured variables (i.e., unmeasured confounders) on the states and actions. If ignoring them, a biased estimation of the policy would be entailed. To break up this confounding gap, in this paper, we take the best of the strong power of instrumental variables (IV) and propose a Confounded Causal Imitation Learning (C2L) model. This model accommodates confounders that influence actions across multiple timesteps, rather than being restricted to immediate temporal dependencies. We develop a two-stage imitation learning framework for valid IV identification and policy optimization. In particular, in the first stage, we construct a testing criterion based on the defined pseudo-variable, with which we achieve identifying a valid IV for the C2L models. Such a criterion entails the sufficient and necessary identifiability conditions for IV validity. In the second stage, with the identified IV, we propose two candidate policy learning approaches: one is based on a simulator, while the other is offline. Extensive experiments verified the effectiveness of identifying the valid IV as well as learning the policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。