arXiv:2509.09208cs.LGcs.AI2025-09IJCAI被引 4

提出新算法IP3O,让强化学习更安全稳定地逼近约束边界。

Incentivizing Safer Actions in Policy Optimization for Constrained Reinforcement Learning

  • 用渐增惩罚机制替代固定奖励,提前激励安全动作
  • 在基准环境上优于现有最先进安全强化学习方法
  • 理论证明算法最坏情况下的最优性误差有界,适合高安全要求场景

约束强化学习旨在最大化回报的同时满足预设约束,以体现领域特定的安全要求。在连续控制任务中,学习代理需决策系统动作,如何平衡奖励最大化与约束满足仍是挑战。现有策略优化方法在约束边界附近常出现不稳定,导致训练性能下降。为此,我们引入一种自适应激励机制,结合奖励结构,在接近约束边界前就促使智能体保持安全行为。基于此,提出增量惩罚近端策略优化(IP3O)算法,通过逐步增加惩罚来稳定训练动态。在基准环境上的实证评估表明,该算法在性能上优于当前最先进的安全强化学习方法。此外,我们还推导出算法实现最优性的最坏情况误差上界,提供理论保障。

原文摘要 · Abstract (English)

Constrained Reinforcement Learning (RL) aims to maximize the return while adhering to predefined constraint limits, which represent domain-specific safety requirements. In continuous control settings, where learning agents govern system actions, balancing the trade-off between reward maximization and constraint satisfaction remains a significant challenge. Policy optimization methods often exhibit instability near constraint boundaries, resulting in suboptimal training performance. To address this issue, we introduce a novel approach that integrates an adaptive incentive mechanism in addition to the reward structure to stay within the constraint bound before approaching the constraint boundary. Building on this insight, we propose Incrementally Penalized Proximal Policy Optimization (IP3O), a practical algorithm that enforces a progressively increasing penalty to stabilize training dynamics. Through empirical evaluation on benchmark environments, we demonstrate the efficacy of IP3O compared to the performance of state-of-the-art Safe RL algorithms. Furthermore, we provide theoretical guarantees by deriving a bound on the worst-case error of the optimality achieved by our algorithm.

强化学习安全控制策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。