通过双向策略协作提升强化学习的安全性与稳定性
Regularized Reward-Punishment Reinforcement Learning

- 用动态先验机制让奖励与惩罚策略直接互为参考
- 在网格与机器人导航任务中实现更安全且稳定的训练
- 适合需要多目标动机协调的复杂强化学习场景
我们提出KL耦合策略正则化(KCPR),一种用于奖励-惩罚强化学习(RPRL)的策略协同框架。基于KCPR,推导出KL耦合软最优(KCSO)并实现其深度版本klDMP。与现有方法中独立优化奖励与惩罚策略不同,KCPR通过将每个策略视为另一个的动态先验,实现策略间的直接交互。KCSO生成耦合的软最优策略和KL正则化的贝尔曼算子,使奖励与惩罚信息共同影响价值传播。为提升学习稳定性,引入同伴先验平滑机制,并评估了分别存储奖励与惩罚经验的回放缓冲设计。在网格世界和Gazebo机器人导航任务中的实验表明,klDMP在保持与DQN、SQL和softDMP相当的任务性能的同时,显著提升了安全性与学习稳定性。结果表明,策略级协同是整合多重行为目标的有效机制,可作为具有交互动机过程的强化学习系统的设计原则。
原文摘要 · Abstract (English)
We propose KL-Coupled Policy Regularization (KCPR), a policy coordination framework for Reward-Punishment Reinforcement Learning (RPRL). Based on KCPR, we derive KL-Coupled Soft Optimality (KCSO) and develop its deep realization, klDMP. Unlike existing RPRL approaches that optimize reward-seeking and punishment-related policies largely independently, KCPR enables direct interactions between companion policies by treating each as a dynamically learned prior for the other. KCSO yields coupled soft-optimal policies and KL-regularized Bellman operators, allowing reward and punishment information to jointly influence value propagation. To improve learning stability, we introduce a companion-prior softening mechanism and evaluate separate replay-buffer designs for balancing reward- and punishment-related experience. Experiments in grid-world and Gazebo robotic navigation tasks demonstrate that klDMP improves safety and learning stability while maintaining competitive task performance compared with DQN, SQL and softDMP. These results suggest that policy-level coordination provides an effective mechanism for integrating multiple behavioral objectives and may serve as a useful design principle for reinforcement learning systems with interacting motivational processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。