arXiv:2410.11250cs.AIcs.LG2024-10被引 2

改进强化学习在连续空间中的表现,提升模型鲁棒性。

Learning Agents With Prioritization and Parameter Noise in Continuous State and Action Space

  • 结合DQN与DDPG思想,引入优先级机制优化学习过程。
  • 训练中加入参数噪声,使模型在连续状态动作空间上性能更优。
  • 适合研究机器人控制、自动驾驶等连续决策问题的学者参考。

在众多强化学习变体中,状态和动作空间均为连续的问题尤为重要——自主机器人、自动驾驶车辆、最优控制等均属于此类,天然适配基于强化学习的算法。本文提出一种结合前沿方法(如Deep Q-learning, DQN 和 Deep Deterministic Policy Gradient, DDPG)的优先级形式,显著优于此前连续状态-动作空间问题的基准结果。实验还引入训练期间的参数噪声,进一步提升了深度强化学习模型的鲁棒性,性能超越已有成果。这些进展对连续空间强化学习问题具有重要价值。

原文摘要 · Abstract (English)

Among the many variants of RL, an important class of problems is where the state and action spaces are continuous -- autonomous robots, autonomous vehicles, optimal control are all examples of such problems that can lend themselves naturally to reinforcement based algorithms, and have continuous state and action spaces. In this paper, we introduce a prioritized form of a combination of state-of-the-art approaches such as Deep Q-learning (DQN) and Deep Deterministic Policy Gradient (DDPG) to outperform the earlier results for continuous state and action space problems. Our experiments also involve the use of parameter noise during training resulting in more robust deep RL models outperforming the earlier results significantly. We believe these results are a valuable addition for continuous state and action space problems.

强化学习连续控制鲁棒性深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。