arXiv:2409.10218cs.LG2024-09被引 10

通过剪枝与模型检测结合,让强化学习更安全可解释

Safety-Oriented Pruning and Interpretation of Reinforcement Learning Policies

  • 用模型检测精确分析剪枝对安全性的具体影响
  • 保持剪枝后策略的安全性,且能量化连接重要性
  • 适合关注安全可控的RL应用开发者

神经网络剪枝虽能简化模型,但可能误删强化学习(RL)策略中关键参数。本文提出一种可解释的强化学习方法VERINTER,将神经网络剪枝与模型检测相结合,确保可解释的强化学习安全性。VERINTER通过分析安全度量的变化,精确量化剪枝的影响及神经连接对复杂安全属性的作用。该方法在多个强化学习场景中均验证有效,既能保持剪枝后策略的安全性,又增强了对其安全动态的理解。

原文摘要 · Abstract (English)

Pruning neural networks (NNs) can streamline them but risks removing vital parameters from safe reinforcement learning (RL) policies. We introduce an interpretable RL method called VERINTER, which combines NN pruning with model checking to ensure interpretable RL safety. VERINTER exactly quantifies the effects of pruning and the impact of neural connections on complex safety properties by analyzing changes in safety measurements. This method maintains safety in pruned RL policies and enhances understanding of their safety dynamics, which has proven effective in multiple RL settings.

强化学习模型剪枝可解释性安全约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。