arXiv:2502.07656cs.LGcs.AI2025-02被引 1

解决专家可见但模仿者不可见的隐藏干扰问题,提升强化学习模仿性能。

Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding

  • 用轨迹历史作工具变量,将因果模仿学习转为条件矩约束问题。
  • 在Mujoco任务上,新方法比现有基线降低模仿误差,性能更优。
  • 适合处理存在隐蔽干扰的复杂模仿学习场景,如机器人控制。

我们提出一种通用的因果模仿学习框架,用于处理隐藏混淆因素,涵盖多种现有设置。该框架考虑两类隐藏混淆:(a) 专家可观测但模仿者不可观测的变量;(b) 对双方均隐藏的混淆噪声。通过利用轨迹历史作为工具变量,我们将该框架下的因果模仿学习重构为条件矩约束(CMR)问题。我们提出DML-IL算法,通过工具变量回归求解此CMR问题,并给出了模仿差距的上界。在连续状态动作环境(包括Mujoco任务)上的实证评估表明,DML-IL优于现有因果模仿学习基线。

原文摘要 · Abstract (English)

We propose a general framework for causal Imitation Learning (IL) with hidden confounders, which subsumes several existing settings. Our framework accounts for two types of hidden confounders: (a) variables observed by the expert but not by the imitator, and (b) confounding noise hidden from both. By leveraging trajectory histories as instruments, we reformulate causal IL in our framework into a Conditional Moment Restriction (CMR) problem. We propose DML-IL, an algorithm that solves this CMR problem via instrumental variable regression, and upper bound its imitation gap. Empirical evaluation on continuous state-action environments, including Mujoco tasks, demonstrates that DML-IL outperforms existing causal IL baselines.

因果学习模仿学习工具变量强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。