arXiv:2510.18478cs.LG2025-10中稿 · to AAMAS '26

通过不确定性调制减少安全批评家的过度保守,提升强化学习安全性与性能平衡。

Safe But Not Sorry: Reducing Over-Conservatism in Safety Critics via Uncertainty-Aware Modulation

  • 在不确定高风险区域加强保守性,安全区域保持梯度清晰
  • 安全违规降低约40%,成本梯度预测误差减少约83%
  • 适合需高安全性的强化学习部署场景

确保强化学习(RL)代理在真实系统中的安全探索至关重要。现有方法难以平衡:严格安全约束会严重损害任务性能,而过度追求奖励则常导致安全约束被违反,产生平坦的成本景观,削弱梯度并阻碍策略优化。我们提出不确定安全批评家(USC),将不确定性感知的调制与优化融入批评家训练中。通过在不确定和高成本区域集中保守性,同时在安全区域保持清晰梯度,USC实现了有效的奖励-安全权衡。大量实验表明,USC使安全违规减少约40%,保持或超过竞争性奖励水平,且预测与真实成本梯度误差降低约83%,打破了安全与性能之间的固有权衡,为可扩展的安全强化学习铺平道路。

原文摘要 · Abstract (English)

Ensuring the safe exploration of reinforcement learning (RL) agents is critical for deployment in real-world systems. Yet existing approaches struggle to strike the right balance: methods that tightly enforce safety often cripple task performance, while those that prioritize reward leave safety constraints frequently violated, producing diffuse cost landscapes that flatten gradients and stall policy improvement. We introduce the Uncertain Safety Critic (USC), a novel approach that integrates uncertainty-aware modulation and refinement into critic training. By concentrating conservatism in uncertain and costly regions while preserving sharp gradients in safe areas, USC enables policies to achieve effective reward-safety trade-offs. Extensive experiments show that USC reduces safety violations by approximately 40% while maintaining competitive or higher rewards, and reduces the error between predicted and true cost gradients by approximately 83%, breaking the prevailing trade-off between safety and performance and paving the way for scalable safe RL.

强化学习安全控制不确定性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。