arXiv:2504.19473cs.LGcs.RO2025-04被引 1

用自适应控制李雅普诺夫函数提升强化学习的安全性与稳定性

Stability Enhancement in Reinforcement Learning via Adaptive Control Lyapunov Function

  • 设计任务相关的控制李雅普诺夫函数,确保安全与性能平衡
  • 动态调整约束条件,增强对未建模动态的鲁棒性
  • 在保证安全的同时提升控制输入平滑性,适合真实系统部署

强化学习在控制任务中展现出潜力,但在实际应用中因缺乏学习过程中的安全保证而面临挑战。现有方法难以确保安全探索,易导致系统故障,限制其在仿真环境外的应用。传统奖励塑造和约束策略优化在初始学习阶段无法保证安全,基于模型的方法如使用控制李雅普诺夫函数(CLF)或控制屏障函数(CBF)则可能抑制有效探索。本文提出软演员-评论家结合控制李雅普诺夫函数(SAC-CLF)框架,通过三项创新实现:(1) 针对任务的CLF设计方法,实现安全与最优性能;(2) 动态调整约束以应对未建模动态;(3) 提升控制输入平滑性并维持安全。在典型非线性系统与卫星姿态控制实验中验证了该方法的有效性。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has shown promise in control tasks but faces significant challenges in real-world applications, primarily due to the absence of safety guarantees during the learning process. Existing methods often struggle with ensuring safe exploration, leading to potential system failures and restricting applications primarily to simulated environments. Traditional approaches such as reward shaping and constrained policy optimization can fail to guarantee safety during initial learning stages, while model-based methods using Control Lyapunov Functions (CLFs) or Control Barrier Functions (CBFs) may hinder efficient exploration and performance. To address these limitations, this paper introduces Soft Actor-Critic with Control Lyapunov Function (SAC-CLF), a framework that enhances stability and safety through three key innovations: (1) a task-specific CLF design method for safe and optimal performance; (2) dynamic adjustment of constraints to maintain robustness under unmodeled dynamics; and (3) improved control input smoothness while ensuring safety. Experimental results on a classical nonlinear system and satellite attitude control demonstrate the effectiveness of SAC-CLF in overcoming the shortcomings of existing methods.

强化学习安全控制李雅普诺夫函数稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。