用因果表示提升离线元强化学习的泛化能力
CausalCOMRL: Context-Based Offline Meta-Reinforcement Learning with Causal Representation
- 通过因果表示学习挖掘任务组件间的真实关系
- 在多个基准上性能优于现有方法
- 适合需要强泛化能力的复杂任务场景
基于上下文的离线元强化学习(OMRL)方法利用预收集的数据集,通过任务表征指导策略学习,取得了显著进展。然而,现有方法常引入虚假相关性,即因混杂因素导致任务成分间错误关联,当测试任务中的混杂因素与训练不同时,会降低策略性能。为此,我们提出CausalCOMRL,一种融合因果表示学习的上下文型离线元强化学习方法。该方法揭示任务成分间的因果关系,并将其融入任务表征,提升强化学习代理的泛化能力。进一步通过互信息优化和对比学习增强不同任务表征的区分度。基于这些因果表征,采用SAC在元强化学习基准上优化策略。实验表明,CausalCOMRL在多数基准上表现优于其他方法。
原文摘要 · Abstract (English)
Context-based offline meta-reinforcement learning (OMRL) methods have achieved appealing success by leveraging pre-collected offline datasets to develop task representations that guide policy learning. However, current context-based OMRL methods often introduce spurious correlations, where task components are incorrectly correlated due to confounders. These correlations can degrade policy performance when the confounders in the test task differ from those in the training task. To address this problem, we propose CausalCOMRL, a context-based OMRL method that integrates causal representation learning. This approach uncovers causal relationships among the task components and incorporates the causal relationships into task representations, enhancing the generalizability of RL agents. We further improve the distinction of task representations from different tasks by using mutual information optimization and contrastive learning. Utilizing these causal task representations, we employ SAC to optimize policies on meta-RL benchmarks. Experimental results show that CausalCOMRL achieves better performance than other methods on most benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。