将因果推理与强化学习结合,揭示二者在反事实分析上的深层联系。
An Introduction to Causal Reinforcement Learning

- 用结构因果模型统一建模强化学习环境中的因果机制
- 提出反事实学习、模仿学习等新范式,拓展学习维度
- 适合对因果推断和强化学习交叉研究感兴趣的读者
因果推断提供了一套原则与工具,使我们能够结合数据与环境知识,回答反事实问题——即在现实不同情况下会发生什么,即使没有该情境的实际数据。强化学习则通过试错方式学习优化特定目标(如奖励、遗憾)的策略。这两门学科长期独立发展,几乎无交集。我们指出,它们均作用于同一核心要素——反事实关系,具有本质关联。当这种联系被明确识别并数学化时,新的学习机会涌现。我们发现,任何强化学习环境都可分解为一系列具有不同因果不变性的自主机制,可用结构因果模型简洁建模;标准强化学习设置隐含此类模型。该形式化使得在线学习、离线策略学习与因果计算学习得以统一处理。但这些模式并不完整:我们引入并探讨了广义策略学习、干预选择、模仿学习与反事实学习等自然且普遍的学习场景。这些任务拓宽了反事实学习的视野,提示因果推断与强化学习并行研究的巨大潜力,我们称之为因果强化学习(CRL)。
原文摘要 · Abstract (English)
Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i.e., what would have happened had reality been different, even when no data of this unrealized reality is currently available. Reinforcement learning provides methods to learn a policy that optimizes a specific measure (e.g., reward, regret) when the agent is deployed in an environment and pursues an exploratory, trial-and-error approach. These two disciplines have evolved independently and with virtually no interaction between them. We note that they operate over different aspects of the same building block, counterfactual relations, which makes them umbilically connected. Based on these observations, novel learning opportunities arise when this connection is explicitly acknowledged and mathematized. To realize this potential, we note that any environment where the RL agent is deployed can be decomposed as a collection of autonomous mechanisms with different causal invariances, parsimoniously modeled as a structural causal model; any standard RL setting implicitly encodes such a model. This formalization allows us to put under a unifying treatment different modes of learning, including online, off-policy, and causal calculus learning, which appear unrelated in the literature. However, these modalities are not exhaustive: we introduce several natural and pervasive classes of learning settings that entail novel dimensions of analysis. Specifically, we introduce and discuss through causal lenses generalized policy learning, where to intervene, imitation learning, and counterfactual learning. These tasks lead to a broader view of counterfactual learning and suggest great potential for studying causal inference and reinforcement learning side by side, which we call causal reinforcement learning (CRL).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。