新算法让智能体在延迟反馈下仍能高效决策。
Delay-Empowered Causal Hierarchical Reinforcement Learning

- 通过因果结构与随机延迟建模,实现对延迟的显式理解。
- 在2D-Minecraft和MiniGrid上显著优于基线方法。
- 适合处理延迟不确定的现实任务,如机器人控制、自动化系统。
许多现实任务涉及延迟效应,即动作结果在不同时间滞后后才显现。现有延迟感知强化学习方法通常依赖状态扩展、延迟分布先验知识或非延迟数据,限制了泛化能力。相比之下,层级强化学习因结构优势更适于处理延迟,但现有方法仅适用于固定延迟。为此,我们提出延迟赋能因果层级强化学习(DECHRL),显式建模状态转移的因果结构及其关联的随机延迟分布,并将其融入延迟感知的赋能目标,驱动智能体主动探索高可控状态,从而提升在时间不确定性下的表现。我们在修改后的2D-Minecraft和MiniGrid环境中评估了DECHRL,实验结果表明其能有效建模时间延迟,在时间不确定性下的决策性能显著优于基线方法。
原文摘要 · Abstract (English)
Many real-world tasks involve delayed effects, where the outcomes of actions emerge after varying time lags. Existing delay-aware reinforcement learning methods often rely on state augmentation, prior knowledge of delay distributions, or access to non-delayed data, limiting their generalization. Hierarchical reinforcement learning, by contrast, inherently offers advantages in handling delays due to its hierarchical structure, yet existing methods are restricted to fixed delays. To address these limitations, we propose Delay-Empowered Causal Hierarchical Reinforcement Learning (DECHRL). DECHRL explicitly models both the causal structure of state transitions and their associated stochastic delay distributions. These are then incorporated into a delay-aware empowerment objective that drives proactive exploration toward highly controllable states, thereby improving performance under temporal uncertainty. We evaluate DECHRL in modified 2D-Minecraft and MiniGrid environments featuring stochastic delays. Experimental results show that DECHRL effectively models temporal delays and significantly outperforms baselines in decision-making under temporal uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。