arXiv:2507.18992cs.LGcs.AI2025-07被引 1

提出保守代理,解决随机延迟环境下的强化学习难题

Reinforcement Learning via Conservative Agent for Environments with Random Delays

  • 将随机延迟环境转化为等效常数延迟环境,保持算法结构不变
  • 在连续控制任务中显著提升最终性能与采样效率
  • 适合真实场景中反馈延迟不可预测的强化学习应用

现实世界的强化学习应用常受环境反馈延迟影响,破坏马尔可夫性并带来重大挑战。尽管已有众多针对恒定延迟的补偿方法,但因随机延迟的内在变异性与不可预测性,相关研究仍属空白。本文提出一种简单而鲁棒的决策代理——保守代理,将随机延迟环境重构为等效的常数延迟环境。该转换使任意先进常数延迟方法可直接拓展至随机延迟环境,无需修改算法结构且不损失性能。我们在连续控制任务上评估基于保守代理的算法,实验结果表明其在渐近性能和样本效率方面均显著优于现有基线方法。

原文摘要 · Abstract (English)

Real-world reinforcement learning applications are often hindered by delayed feedback from environments, which violates the Markov assumption and introduces significant challenges. Although numerous delay-compensating methods have been proposed for environments with constant delays, environments with random delays remain largely unexplored due to their inherent variability and unpredictability. In this study, we propose a simple yet robust agent for decision-making under random delays, termed the conservative agent, which reformulates the random-delay environment into its constant-delay equivalent. This transformation enables any state-of-the-art constant-delay method to be directly extended to the random-delay environments without modifying the algorithmic structure or sacrificing performance. We evaluate the conservative agent-based algorithm on continuous control tasks, and empirical results demonstrate that it significantly outperforms existing baseline algorithms in terms of asymptotic performance and sample efficiency.

强化学习随机延迟决策代理连续控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。