arXiv:2508.01883cs.LG2025-08AAAI被引 3

提出提前惩罚机制,让强化学习更安全稳定

Proactive Constrained Policy Optimization with Preemptive Penalty

  • 在策略接近约束边界时提前施加惩罚,避免违规
  • 实验显示该方法显著提升优化稳定性,减少震荡
  • 适合对安全性要求高的机器人控制等场景

安全强化学习常面临约束违反和不稳定性问题,传统方法多采用拉格朗日法,属于事后修正,易导致振荡和超调。为此,本文提出主动约束策略优化(PCPO),引入前瞻惩罚机制:当策略逼近约束边界时,通过在目标函数中加入障碍项施加成本。同时设计一种仅在接近边界时激活的内在奖励,引导边界感知探索。理论分析给出了对偶间隙和更新性能的上下界,揭示收敛特性。为提升优化效果,采用策略迭代框架。实验表明,PCPO具有显著稳定性,为约束下的策略优化提供了稳健解决方案,对后续研究与实际应用有重要意义。

原文摘要 · Abstract (English)

Safe Reinforcement Learning (RL) often faces significant issues such as constraint violations and instability, necessitating the use of constrained policy optimization, which seeks optimal policies while ensuring adherence to specific constraints like safety. Typically, constrained optimization problems are addressed by the Lagrangian method, a post-violation remedial approach that may result in oscillations and overshoots. Motivated by this, we propose a novel method named Proactive Constrained Policy Optimization (PCPO) that incorporates a preemptive penalty mechanism. This mechanism integrates barrier items into the objective function as the policy nears the boundary, imposing a cost. Meanwhile, we introduce a constraint-aware intrinsic reward to guide boundary-aware exploration, which is activated only when the policy approaches the constraint boundary. We establish theoretical upper and lower bounds for the duality gap and the performance of the PCPO update, shedding light on the method's convergence characteristics. Additionally, to enhance the optimization performance, we adopt a policy iteration approach. An interesting finding is that PCPO demonstrates significant stability in experiments. Experimental results indicate that the PCPO framework provides a robust solution for policy optimization under constraints, with important implications for future research and practical applications.

强化学习安全控制约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。