让智能体在延迟环境下仍能跨任务快速迁移,靠的是隐式因果图建模。
Transferable Delay-Aware Reinforcement Learning via Implicit Causal Graph Modeling

- 用隐式因果图建模动作与状态间的动态依赖关系。
- 在带随机延迟的连续控制任务中超越基线方法。
- 学习到的结构化表征可有效迁移到新任务,加速适应。
随机延迟会削弱动作与后续状态反馈之间的时序对应关系,使智能体难以识别动作效应的真实传播过程。在跨任务场景中,任务目标与奖励设计的变化进一步降低了已有任务知识的可复用性。为此,本文提出一种基于隐式因果图建模的可迁移延迟感知强化学习方法。该方法采用场-节点编码器将高维观测映射为具有节点语义的潜在状态,并通过消息传递机制刻画节点间的动态因果依赖,从而学习可迁移的结构化表征与环境动态知识。在此基础上,引入想象驱动的行为学习与规划,在潜在空间优化策略,实现跨任务知识迁移与快速适应。实验结果表明,所提方法在带有随机延迟的DMC连续控制任务上优于基线方法;跨任务迁移实验进一步证明,学习到的结构化表征与动态知识可有效迁移至新任务,显著加速策略适应。
原文摘要 · Abstract (English)
Random delays weaken the temporal correspondence between actions and subsequent state feedback, making it difficult for agents to identify the true propagation process of action effects. In cross-task scenarios, changes in task objectives and reward formulations further reduce the reusability of previously acquired task knowledge. To address this problem, this paper proposes a transferable delay-aware reinforcement learning method based on implicit causal graph modeling. The proposed method uses a field-node encoder to represent high-dimensional observations as latent states with node-level semantics, and employs a message-passing mechanism to characterize dynamic causal dependencies among nodes, thereby learning transferable structured representations and environment dynamics knowledge. On this basis, imagination-driven behavior learning and planning are incorporated to optimize policies in the latent space, enabling cross-task knowledge transfer and rapid adaptation. Experimental results show that the proposed method outperforms baseline methods on DMC continuous control tasks with random delays. Cross-task transfer experiments further demonstrate that the learned structured representations and dynamics knowledge can be effectively transferred to new tasks and significantly accelerate policy adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。