arXiv:2510.23424cs.AI2025-10

让强化学习模型学会因果推理,摆脱错误关联

Causal Deep Q Network

  • 用因果公式估计动作真实效果,避免学习虚假关联
  • 在标准环境上性能超越传统DQN,提升问题解决能力
  • 适合想提升智能体决策可靠性的研究者使用

深度Q网络(DQN)在强化学习任务中表现卓越,但其依赖关联学习常导致习得虚假相关性,限制了问题求解能力。本文提出将因果原则融入DQN,利用PEACE(概率易变因果效应)公式估算因果效应。通过在训练中引入因果推理,新框架增强了DQN对环境潜在因果结构的理解,有效缓解混淆因素和虚假相关性的影响。实验表明,融合因果能力的DQN显著提升问题求解能力,且不损害原有性能。在标准基准环境上的结果验证了该方法的有效性。本工作为基于严谨因果推断提升深度强化学习智能体能力提供了可行路径。

原文摘要 · Abstract (English)

Deep Q Networks (DQN) have shown remarkable success in various reinforcement learning tasks. However, their reliance on associative learning often leads to the acquisition of spurious correlations, hindering their problem-solving capabilities. In this paper, we introduce a novel approach to integrate causal principles into DQNs, leveraging the PEACE (Probabilistic Easy vAriational Causal Effect) formula for estimating causal effects. By incorporating causal reasoning during training, our proposed framework enhances the DQN's understanding of the underlying causal structure of the environment, thereby mitigating the influence of confounding factors and spurious correlations. We demonstrate that integrating DQNs with causal capabilities significantly enhances their problem-solving capabilities without compromising performance. Experimental results on standard benchmark environments showcase that our approach outperforms conventional DQNs, highlighting the effectiveness of causal reasoning in reinforcement learning. Overall, our work presents a promising avenue for advancing the capabilities of deep reinforcement learning agents through principled causal inference.

强化学习因果推理DQN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。