通过神经元激活图分析,揭示安全强化学习策略的决策机制。
Co-Activation Graph Analysis of Safety-Verified and Explainable Deep Reinforcement Learning Policies
- 结合模型检测与神经元激活模式分析,解析策略内部逻辑。
- 可识别安全决策背后的神经元协同激活模式。
- 适合关注AI安全与可解释性的研究人员使用。
深度强化学习策略可能表现出不安全行为且难以解释。为应对这些挑战,本文将强化学习策略模型检测(用于判断策略是否具有不安全行为)与共激活图分析(通过分析神经元激活模式揭示神经网络内部运作)相结合,深入理解安全强化学习策略的序列决策过程。该方法使我们能够解释策略在安全决策中的内部工作机制。我们在多个实验中验证了该方法的有效性。
原文摘要 · Abstract (English)
Deep reinforcement learning (RL) policies can demonstrate unsafe behaviors and are challenging to interpret. To address these challenges, we combine RL policy model checking--a technique for determining whether RL policies exhibit unsafe behaviors--with co-activation graph analysis--a method that maps neural network inner workings by analyzing neuron activation patterns--to gain insight into the safe RL policy's sequential decision-making. This combination lets us interpret the RL policy's inner workings for safe decision-making. We demonstrate its applicability in various experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。