arXiv:2412.13184cs.LGcs.AI2024-12中稿 · AAAI被引 3

用分位数梯度更新提升强化学习安全性,兼顾高回报与严格约束。

Tilted Quantile Gradient Updates for Quantile-Constrained Reinforcement Learning

  • 直接采样估计分位数梯度,避免期望形式近似
  • 在3个基准环境中100%满足安全约束,回报高于现有方法
  • 适合对安全要求严苛的机器人控制、自动驾驶场景

安全强化学习是实现奖励最大化策略并保证安全性的流行范式。以往工作多采用期望形式表达安全约束,虽易实现但难以在高概率下维持安全。为此,本文转向分位数约束强化学习,无需期望形式近似即可实现更高水平的安全性。通过采样直接估计分位数梯度,并给出收敛性理论证明。进一步提出倾斜梯度更新策略,补偿分布密度的非对称性,显著提升回报表现。实验表明,所提模型在多个基准任务中100%满足分位数约束,同时超越现有最先进方法的回报表现。

原文摘要 · Abstract (English)

Safe reinforcement learning (RL) is a popular and versatile paradigm to learn reward-maximizing policies with safety guarantees. Previous works tend to express the safety constraints in an expectation form due to the ease of implementation, but this turns out to be ineffective in maintaining safety constraints with high probability. To this end, we move to the quantile-constrained RL that enables a higher level of safety without any expectation-form approximations. We directly estimate the quantile gradients through sampling and provide the theoretical proofs of convergence. Then a tilted update strategy for quantile gradients is implemented to compensate the asymmetric distributional density, with a direct benefit of return performance. Experiments demonstrate that the proposed model fully meets safety requirements (quantile constraints) while outperforming the state-of-the-art benchmarks with higher return.

强化学习安全控制分位数优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。