通过扰动动作学习因果简化模型,解释强化学习策略的成败原因
Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies
- 用随机扰动动作观测奖励变化,构建高阶因果模型
- 在摆杆控制和机器人乒乓球任务中发现策略偏差与失败模式
- 保证干预一致性,解释结果真实反映因果关系
强化学习策略为何成功或失败?由于智能体与环境交互的复杂性和高维性,这一问题极具挑战。本文从因果视角出发,将状态、动作和奖励视为低层因果模型中的变量。通过在执行过程中对策略动作施加随机扰动,并观察其对累积奖励的影响,我们学习一个简化的高层因果模型来解释这些关系。为此,我们提出了非线性因果模型降维框架,确保近似干预一致性——即简化后的高层模型对干预的响应方式与原始复杂系统相似。我们证明,在一类非线性因果模型中,存在唯一解可实现精确干预一致性,确保所学解释反映真实的因果模式。在合成因果模型及实际强化学习任务(包括摆杆控制与机器人乒乓球)上的实验表明,该方法能有效揭示训练后策略中的重要行为模式、偏差与失败机制。
原文摘要 · Abstract (English)
Why do reinforcement learning (RL) policies fail or succeed? This is a challenging question due to the complex, high-dimensional nature of agent-environment interactions. In this work, we take a causal perspective on explaining the behavior of RL policies by viewing the states, actions, and rewards as variables in a low-level causal model. We introduce random perturbations to policy actions during execution and observe their effects on the cumulative reward, learning a simplified high-level causal model that explains these relationships. To this end, we develop a nonlinear Causal Model Reduction framework that ensures approximate interventional consistency, meaning the simplified high-level model responds to interventions in a similar way as the original complex system. We prove that for a class of nonlinear causal models, there exists a unique solution that achieves exact interventional consistency, ensuring learned explanations reflect meaningful causal patterns. Experiments on both synthetic causal models and practical RL tasks-including pendulum control and robot table tennis-demonstrate that our approach can uncover important behavioral patterns, biases, and failure modes in trained RL policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。