用动力系统理论分析强化学习安全,揭示隐藏的危险边界与失效模式。
A Dynamical Systems Framework for Reinforcement Learning Safety and Robustness Verification
- 将智能体与环境视为离散动力系统,用FTLE识别行为骨架
- 发现排斥型结构是安全屏障,吸引型结构暴露潜在陷阱状态
- 提出量化指标评估安全裕度,适合验证高风险决策系统
强化学习在安全关键系统中的应用受限于缺乏对策略鲁棒性与安全性的形式化验证方法。本文提出一种新框架,将强化学习智能体及其环境建模为离散时间自治动力系统。基于动力系统理论,利用有限时间李雅普诺夫指数(FTLE)识别并可视化拉格朗日相干结构(LCS),这些结构构成系统行为的隐藏‘骨架’。我们证明排斥型LCS可作为不安全区域周围的防护屏障,而吸引型LCS则揭示系统的收敛特性及潜在故障模式,如非预期的‘陷阱’状态。为进一步实现定量分析,引入一组量化指标:平均边界排斥度(MBR)、综合虚假吸引子强度(ASAS)和时序感知虚假吸引子强度(TASAS),用于正式衡量策略的安全边际与鲁棒性。此外,本文还提供局部稳定性保证推导方法,并扩展分析以处理模型不确定性。在离散与连续控制环境中实验表明,该框架能全面且可解释地评估策略行为,成功识别出仅凭奖励表现看似成功但实际存在严重缺陷的策略。
原文摘要 · Abstract (English)
The application of reinforcement learning to safety-critical systems is limited by the lack of formal methods for verifying the robustness and safety of learned policies. This paper introduces a novel framework that addresses this gap by analyzing the combination of an RL agent and its environment as a discrete-time autonomous dynamical system. By leveraging tools from dynamical systems theory, specifically the Finite-Time Lyapunov Exponent (FTLE), we identify and visualize Lagrangian Coherent Structures (LCS) that act as the hidden "skeleton" governing the system's behavior. We demonstrate that repelling LCS function as safety barriers around unsafe regions, while attracting LCS reveal the system's convergence properties and potential failure modes, such as unintended "trap" states. To move beyond qualitative visualization, we introduce a suite of quantitative metrics, Mean Boundary Repulsion (MBR), Aggregated Spurious Attractor Strength (ASAS), and Temporally-Aware Spurious Attractor Strength (TASAS), to formally measure a policy's safety margin and robustness. We further provide a method for deriving local stability guarantees and extend the analysis to handle model uncertainty. Through experiments in both discrete and continuous control environments, we show that this framework provides a comprehensive and interpretable assessment of policy behavior, successfully identifying critical flaws in policies that appear successful based on reward alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。