arXiv:2602.09207cs.LGcs.AI2026-02

让扩散策略学会识别真正影响回报的因果动作,提升强化学习效果。

CausalGDP: Causality-Guided Diffusion Policies for Reinforcement Learning

  • 引入因果建模,用真实因果关系指导动作生成
  • 在高维控制任务中表现优于现有扩散策略
  • 适合需要可解释性与鲁棒性的复杂决策场景

强化学习在序列决策问题中取得显著进展,基于扩散的策略通过建模高维动作分布进一步提升性能。然而,现有扩散策略主要依赖统计关联,未显式考虑状态、动作与奖励间的因果关系,难以识别真正带来高回报的动作成分。本文提出因果引导的扩散策略(CausalGDP),通过离线数据学习基础扩散策略和初始因果动态模型,捕捉三者间的因果依赖。在实时交互中,持续更新因果信息作为引导信号,驱动扩散过程向那些因果上影响未来状态与奖励的动作聚焦。通过显式建模因果而非仅依赖关联,CausalGDP将优化重点集中于真正提升性能的动作组件。实验表明,该方法在复杂高维控制任务中持续优于或媲美当前最先进的扩散及离线强化学习方法。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has achieved remarkable success in a wide range of sequential decision-making problems. Recent diffusion-based policies further improve RL by modeling complex, high-dimensional action distributions. However, existing diffusion policies primarily rely on statistical associations and fail to explicitly account for causal relationships among states, actions, and rewards, limiting their ability to identify which action components truly cause high returns. In this paper, we propose Causality-guided Diffusion Policy (CausalGDP), a unified framework that integrates causal reasoning into diffusion-based RL. CausalGDP first learns a base diffusion policy and an initial causal dynamical model from offline data, capturing causal dependencies among states, actions, and rewards. During real-time interaction, the causal information is continuously updated and incorporated as a guidance signal to steer the diffusion process toward actions that causally influence future states and rewards. By explicitly considering causality beyond association, CausalGDP focuses policy optimization on action components that genuinely drive performance improvements. Experimental results demonstrate that CausalGDP consistently achieves competitive or superior performance over state-of-the-art diffusion-based and offline RL methods, especially in complex, high-dimensional control tasks.

强化学习扩散模型因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。