解决复杂环境下的模仿学习偏差问题,提升长序列任务表现。
Scalable Causal Imitation Learning

- 基于因果调整框架,用滑动窗口简化长期决策的混淆因子处理
- 在长时序连续控制任务中超越专家表现,误差显著降低
- 适合高维状态空间、存在隐藏干扰的现实场景应用
模仿学习可在未知环境中通过专家示范学习策略,但当模仿者与专家观测不匹配且专家示范中存在未观测混杂因素时性能下降。通过顺序π-后门准则识别合适调整集,因果模仿学习(CIL)为从受干扰数据中近似专家策略提供了框架。然而,现有方法如因果行为克隆(Causal BC)和因果生成对抗模仿学习(Causal GAIL)仅适用于短时程、低维场景。在具有长时程和高维状态-动作空间的连续控制任务中,这些方法表现差:Causal BC出现误差累积,Causal GAIL不稳定且样本效率低,顺序π-后门调整亦不切实际。本文提出因果软Q模仿学习(Causal SQIL)和因果逆软Q学习(Causal IQ-Learn),两种离策略因果模仿学习算法,将因果调整框架与先进逆强化学习目标结合。两者基于对顺序π-后门准则的高效近似生成的因果调整状态表示,利用连续控制环境的因果结构,将全时程调整简化为固定大小的滑动窗口。我们在一系列受干扰环境中评估所有方法,结果表明,Causal SQIL和Causal IQ-Learn在长时程任务中显著优于先前的CIL算法,有时甚至超越专家;而所有无因果意识的模仿方法均无法学习到有效行为。
原文摘要 · Abstract (English)
Imitation learning enables learning a policy in an unknown environment with a latent reward signal using expert demonstrations, but it struggles when the imitator's and expert's observations are mismatched and unobserved confounders are present in expert demonstrations. By identifying appropriate adjustment sets via the sequential $π$-backdoor criterion, causal imitation learning (CIL) provides a framework for approximating the expert's policy from confounded data. However, existing CIL methods, Causal Behavioral Cloning (Causal BC) and Causal Generative Adversarial Imitation Learning (Causal GAIL), are designed for short-horizon, low-dimensional settings. When applied to continuous control tasks with long horizons and high-dimensional state-action spaces, these methods exhibit poor performance: Causal BC suffers from compounding errors, Causal GAIL is unstable and sample-inefficient, and sequential $π$-backdoor adjustment becomes impractical. We introduce Causal Soft Q Imitation Learning (SQIL) and Causal Inverse soft-Q Learning (IQ-Learn), two off-policy causal imitation learning algorithms that combine the causal adjustment framework with state-of-the-art inverse reinforcement learning objectives. Both algorithms operate on causally-adjusted state representations produced by an efficient approximation of the sequential $π$-backdoor criterion, exploiting the causal structure of continuous control environments to reduce the full-horizon adjustment to a fixed-size sliding window. We evaluate all methods in a suite of confounded environments and find that Causal SQIL and Causal IQ-Learn substantially outperform prior CIL algorithms on long-horizon tasks, sometimes surpassing the expert, whereas all causally unaware imitation methods fail to learn meaningful behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。