arXiv:2607.05064cs.AI2026-07

用扩散模型建模延迟状态与真实状态的差异,提升延迟反馈下的强化学习性能。

Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization

论文配图:Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization
图 1 · 摘自论文原文
  • 用扩散模型显式建模延迟状态与真实状态的关系
  • 在多种随机延迟下性能优于现有方法,长延迟场景仍有效
  • 适合处理带随机延迟的连续控制任务,如机器人控制

现实世界中的强化学习常因反馈延迟导致性能严重下降。现有方法多通过构建增强状态或预测真实状态来缓解,但忽略了随机马尔可夫决策过程(stochastic MDP)引起的延迟状态与真实状态之间的固有差异。本文理论上证明了该差异的存在,并表明其会导致最优策略退化。为此,提出扩散引导的不确定性感知延迟策略优化方法(DUPO),利用扩散模型建模延迟状态消息与当前状态的关系,基于推断出的差异估计对延迟策略进行加权。在多个具有随机延迟的连续机器人控制任务上的大量实验表明,DUPO持续优于现有方法,即使在长且随机的延迟场景下仍保持有效性。

原文摘要 · Abstract (English)

Reinforcement learning in real world environments often suffers from severe performance degradation due to delayed feedback. Existing approaches typically mitigate performance degradation caused by observation delays by constructing augmented states or predicting the true states. However, these methods often overlook the inherent discrepancy between delayed state and true states induced by stochastic MDP. We theoretically prove the existence of such a discrepancy and show that it leads to the degradation of the optimal policy. To address this challenge, we propose Diffusion Guided Uncertainty Aware Delayed Policy Optimization (DUPO). Our method explicitly models the relationship between delayed state message and the current state using a diffusion model, and leverages the resulting discrepancy estimates to weight delayed policies. Extensive experiments on continuous robotic control tasks with multiple stochastic delays demonstrate that DUPO consistently outperforms existing methods and remains effective even under long and random delay scenarios.

强化学习延迟反馈扩散模型机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。