arXiv:2505.01979cs.LG2025-05被引 3

解决强化学习中延迟效应与虚假相关问题,提升复杂环境决策可靠性。

D3HRL: A Distributed Hierarchical Reinforcement Learning Approach Based on Causal Discovery and Spurious Correlation Detection

  • 通过分布式因果发现建模跨时滞的因果关系
  • 利用条件独立性检验剔除虚假相关,准确识别真实因果链
  • 适合需要长时序推理与高可靠决策的复杂任务

当前层次化强化学习(HRL)算法在长时序序列决策任务中表现优异,但仍面临延迟效应与虚假相关性两大挑战。为此,本文提出一种基于因果发现与虚假相关检测的分布式层次化强化学习方法D3HRL。首先,将延迟效应建模为不同时间跨度间的因果关系,并采用分布式因果发现学习这些关系;其次,通过条件独立性检验消除虚假相关;最后,基于识别出的真实因果关系构建并训练层次化策略。三个步骤迭代执行,逐步探索任务的完整因果链。在2D-MineCraft与MiniGrid上的实验表明,D3HRL对延迟效应具有更强敏感性,能准确识别因果关系,从而在复杂环境中实现更可靠的决策。

原文摘要 · Abstract (English)

Current Hierarchical Reinforcement Learning (HRL) algorithms excel in long-horizon sequential decision-making tasks but still face two challenges: delay effects and spurious correlations. To address them, we propose a causal HRL approach called D3HRL. First, D3HRL models delayed effects as causal relationships across different time spans and employs distributed causal discovery to learn these relationships. Second, it employs conditional independence testing to eliminate spurious correlations. Finally, D3HRL constructs and trains hierarchical policies based on the identified true causal relationships. These three steps are iteratively executed, gradually exploring the complete causal chain of the task. Experiments conducted in 2D-MineCraft and MiniGrid show that D3HRL demonstrates superior sensitivity to delay effects and accurately identifies causal relationships, leading to reliable decision-making in complex environments.

强化学习因果发现层次化决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。