用人类偏好学习安全约束,避免风险低估。
Safe Reinforcement Learning with Preference-based Constraint Inference
- 引入新机制建模偏好,捕捉安全成本的非对称重尾特性
- 在多个环境上实现更优的安全性与奖励平衡表现
- 适合需要低成本标注的安全关键型强化学习应用
安全强化学习是安全关键决策的标准范式,但现实中的安全约束常复杂、主观且难以显式定义。现有约束推断方法依赖严格假设或大量专家示范,不适用于多数实际场景。本文提出基于偏好的约束强化学习(PbCRL),针对主流布拉德利-特瑞模型无法刻画安全成本的非对称重尾特性导致的风险低估问题,引入新的死区机制,理论上证明其可促进重尾成本分布,实现更优约束对齐。同时,设计信噪比损失以鼓励基于成本方差的探索,有助于策略学习。采用两阶段训练策略降低在线标注负担并自适应提升约束满足度。实验表明,PbCRL 在真实安全需求对齐方面优于当前最先进基线,在安全性和奖励上均有提升。该方法为安全强化学习中的约束推断提供了有效路径,具备广泛安全关键应用潜力。
原文摘要 · Abstract (English)
Safe reinforcement learning (RL) is a standard paradigm for safety-critical decision making. However, real-world safety constraints can be complex, subjective, and even hard to explicitly specify. Existing works on constraint inference rely on restrictive assumptions or extensive expert demonstrations, which are not realistic in many real-world applications. How to cheaply and reliably learn these constraints is the major challenge we focus on in this study. While inferring constraints from human preferences offers a data-efficient alternative, we identify popular Bradley-Terry (BT) models fail to capture the asymmetric, heavy-tailed nature of safety costs, resulting in risk underestimation. It is still rare in the literature to understand the impacts of BT models on the downstream policy learning. To address the above knowledge gaps, we propose a novel approach namely Preference-based Constrained Reinforcement Learning (PbCRL). We introduce a novel dead zone mechanism into preference modeling and theoretically prove that it encourages heavy-tailed cost distributions, thereby achieving better constraint alignment. Additionally, we incorporate a Signal-to-Noise Ratio (SNR) loss to encourage exploration by cost variances, which is found to benefit policy learning. Further, two-stage training strategy is deployed to lower online labeling burdens while adaptively enhancing constraint satisfaction. Empirical results demonstrate that PbCRL achieves superior alignment with true safety requirements and outperforms state-of-the-art baselines in terms of safety and reward. Our work explores a promising and effective way for constraint inference in Safe RL, with great potential in various safety-critical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。