arXiv:2502.15512cs.LG2025-02被引 1

通过潜空间分析动作稳定性,让强化学习更可解释、更安全。

SALSA-RL: Stability Analysis in the Latent Space of Actions for Reinforcement Learning

  • 将动作建模为潜空间中的动态变量,用线性系统预测其稳定性。
  • 无需修改原有模型,即可在多个基准环境上评估动作风险。
  • 适合关注RL安全性和可解释性的研究者与工业应用者。

现代深度强化学习方法在处理连续动作空间方面取得了显著进展。然而,现实世界的控制系统,尤其是需要精确可靠性能的场景,往往要求对智能体行为进行事前可解释评估,以识别与环境交互中的安全或故障风险。为此,本文提出SALSA-RL(动作潜空间中的稳定性分析),一种新型强化学习框架,将控制动作建模为在潜空间中随时间动态演变的变量。通过使用预训练的编码器-解码器和状态相关的线性系统,该方法实现了局部稳定性分析,可在动作执行前预测动作范数的瞬时增长。实验表明,SALSA-RL可非侵入式地部署于预训练的RL智能体,评估其动作的局部稳定性,且不损害在多种基准环境中的性能。通过提供动作生成过程的可解释分析,SALSA-RL为强化学习系统的设计、分析与理论理解提供了有力工具。

原文摘要 · Abstract (English)

Modern deep reinforcement learning (DRL) methods have made significant advances in handling continuous action spaces. However, real-world control systems, especially those requiring precise and reliable performance, often demand interpretability in the sense of a-priori assessments of agent behavior to identify safe or failure-prone interactions with environments. To address this limitation, this work proposes SALSA-RL (Stability Analysis in the Latent Space of Actions), a novel RL framework that models control actions as dynamic, time-dependent variables evolving within a latent space. By employing a pre-trained encoder-decoder and a state-dependent linear system, this approach enables interpretability through local stability analysis, where instantaneous growth in action-norms can be predicted before their execution. It is demonstrated that SALSA-RL can be deployed in a non-invasive manner for assessing the local stability of actions from pretrained RL agents without compromising on performance across diverse benchmark environments. By enabling a more interpretable analysis of action generation, SALSA-RL provides a powerful tool for advancing the design, analysis, and theoretical understanding of RL systems.

强化学习可解释性稳定性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。