用熵正则化让强化学习更抗干扰,自动学会安全动作
Viability of Future Actions: Robust Safety in Reinforcement Learning via Entropy Regularization
- 通过熵正则化引导模型多保留未来可用动作
- 在扰动下仍能保持约束满足,且接近最优解
- 适合需要高鲁棒性的机器人控制场景
尽管强化学习(RL)取得诸多进展,但在未知扰动下学习能稳健满足状态约束的策略仍是未解难题。本文从新视角分析无模型RL中熵正则化与约束惩罚的协同作用。实验发现,熵正则化会自然引导学习过程最大化未来可行动作的数量,从而提升对动作噪声的鲁棒性。进一步表明,通过用惩罚松弛严格的安全约束,约束RL问题可被任意逼近地转化为无约束问题,进而使用标准无模型RL求解。该重构方法在保持安全性和最优性的同时,显著增强对扰动的韧性。结果表明,熵正则化与鲁棒性之间的关联是值得深入研究的潜在方向,仅通过简单的奖励设计即可实现强化学习的鲁棒安全。
原文摘要 · Abstract (English)
Despite the many recent advances in reinforcement learning (RL), the question of learning policies that robustly satisfy state constraints under unknown disturbances remains open. In this paper, we offer a new perspective on achieving robust safety by analyzing the interplay between two well-established techniques in model-free RL: entropy regularization, and constraints penalization. We reveal empirically that entropy regularization in constrained RL inherently biases learning toward maximizing the number of future viable actions, thereby promoting constraints satisfaction robust to action noise. Furthermore, we show that by relaxing strict safety constraints through penalties, the constrained RL problem can be approximated arbitrarily closely by an unconstrained one and thus solved using standard model-free RL. This reformulation preserves both safety and optimality while empirically improving resilience to disturbances. Our results indicate that the connection between entropy regularization and robustness is a promising avenue for further empirical and theoretical investigation, as it enables robust safety in RL through simple reward shaping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。