arXiv:2503.16799cs.LGcs.AI2025-03ICLR被引 10

用因果视角设计更有效的强化学习教学课程。

Causally Aligned Curriculum Learning

  • 基于因果图判断哪些任务可作为有效教学环节。
  • 在存在隐藏干扰因素时仍能保持策略一致性。
  • 适用于有复杂隐藏变量的视觉强化学习任务。

强化学习中的高维状态-动作空间常导致“维度灾难”。课程学习通过一系列相关且更易处理的源任务训练智能体,期望共享最优决策规则能加速目标任务的学习。然而,当环境存在未观测到的混杂因素时,该假设往往不成立。本文从因果视角研究课程强化学习,推导出确保最优决策规则不变的充分图结构条件,并提出一种基于目标任务因果知识的高效课程生成算法。实验在具有像素观测的离散与连续混杂任务中验证了该方法的有效性。

原文摘要 · Abstract (English)

A pervasive challenge in Reinforcement Learning (RL) is the "curse of dimensionality" which is the exponential growth in the state-action space when optimizing a high-dimensional target task. The framework of curriculum learning trains the agent in a curriculum composed of a sequence of related and more manageable source tasks. The expectation is that when some optimal decision rules are shared across source tasks and the target task, the agent could more quickly pick up the necessary skills to behave optimally in the environment, thus accelerating the learning process. However, this critical assumption of invariant optimal decision rules does not necessarily hold in many practical applications, specifically when the underlying environment contains unobserved confounders. This paper studies the problem of curriculum RL through causal lenses. We derive a sufficient graphical condition characterizing causally aligned source tasks, i.e., the invariance of optimal decision rules holds. We further develop an efficient algorithm to generate a causally aligned curriculum, provided with qualitative causal knowledge of the target task. Finally, we validate our proposed methodology through experiments in discrete and continuous confounded tasks with pixel observations.

强化学习因果推理课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。