arXiv:2602.04132eess.SYcs.LG2026-02被引 1

用李雅普诺夫函数约束强化学习,让机器人控制更稳定可靠。

LC-SAC: Lyapunov-Constrained Soft Actor-Critic via Koopman Operator Theory for Trajectory Tracking and Stabilization

  • 基于柯普曼算子构建系统误差的线性代理模型,推导出闭式二次型控制李雅普诺夫函数。
  • 将李雅普诺夫约束以条件风险价值形式融入智能体更新,聚焦罕见但严重的不稳定性事件。
  • 在小车和四旋翼等高维系统上验证了硬约束的必要性,奖励塑形会破坏学习稳定性。

强化学习在复杂序列决策问题中取得显著进展,但在安全关键型物理系统中的应用受限于缺乏稳定性保证。标准RL算法优先最大化奖励,常导致振荡或状态发散。本文提出基于柯普曼算子理论的李雅普诺夫约束软动作-评论家(LC-SAC)算法。通过扩展动态模态分解(EDMD)学习误差动力学的线性提升代理模型,并求解离散代数里卡蒂方程(DARE),获得闭式二次型候选控制李雅普诺夫函数(CLF)。该CLF作为拉格朗日惩罚项引入SAC智能体更新中,通过条件风险价值(CVaR)目标聚合最坏情况下的违反概率,集中约束压力于罕见但严重的不稳定事件。进一步引入三项结构改进:提升矩阵前的谱半径归一化、物理意义明确的LQR状态代价、以及强制值函数满足V(0)=0的偏置锚点,使闭式CLF在高维提升模型(如小车与3D四旋翼)中具有良好的适定性。消融实验表明,硬拉格朗日约束至关重要,替换为奖励塑形(Lyap-RS-SAC)会破坏学习并导致四旋翼任务回报崩溃。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has achieved remarkable success in solving complex sequential decision-making problems. However, its application to safety-critical physical systems remains constrained by the lack of stability guarantees. Standard RL algorithms prioritize reward maximization, often yielding policies that may induce oscillations or unbounded state divergence. In this work we propose a Lyapunov-Constrained Soft Actor-Critic (LC-SAC) algorithm using Koopman operator theory. We learn a linear lifted surrogate of the error dynamics via Extended Dynamic Mode Decomposition (EDMD) and solve the Discrete Algebraic Riccati Equation (DARE) to obtain a closed-form quadratic candidate Control Lyapunov Function (CLF). This CLF is incorporated into the SAC actor update as a Lagrangian penalty that aggregates the worst-case tail of violations via a Conditional Value-at-Risk (CVaR) objective, concentrating constraint pressure on rare but severe instability events. We further introduce three structural EDMD refinements spectral-radius normalization of the lifted A-matrix prior to the DARE solve, a physically meaningful LQR state cost, and a value-bias anchor enforcing V(0)=0 that make the closed-form CLF well-posed for higher-dimensional lifted models such as the cartpole and 3D quadrotor. The ablation study shows that a hard Lagrangian constraint is essential, replacing it with reward shaping (Lyap-RS-SAC) destabilizes learning and collapses return on quadrotor tasks.

强化学习控制理论稳定性保障柯普曼算子

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。