arXiv:2504.02565eess.SYcs.LG2025-04被引 5

提出MAD策略,让强化学习控制更稳定且无需系统模型

MAD: A Magnitude And Direction Policy Parametrization for Stability Constrained Reinforcement Learning

  • 用幅值-方向分离结构设计策略,保证非线性系统的闭环稳定性
  • 在未知模型下训练的策略仍能保持稳定,泛化性能媲美普通神经网络
  • 适合对安全性要求高的强化学习应用,如机器人控制

我们提出幅度与方向(MAD)策略,一种用于强化学习的策略参数化方法,可确保非线性动态系统的Lp闭环稳定性。现有基于非线性Youla和系统级综合的方法受限于Lp稳定算子的参数化难题。MAD策略通过将控制输入的幅值由扰动反馈的Lp稳定算子描述,方向则由状态依赖特征通过通用函数逼近器确定,既保留了强化学习成功的关键——状态依赖反馈,又不破坏闭环稳定性。我们进一步分析了在模型失配下的鲁棒稳定性。与现有扰动反馈策略不同,MAD策略引入了兼容无模型强化学习的状反馈成分,在仅假设开环稳定的前提下实现闭环稳定。数值实验表明,使用深度确定性策略梯度(DDPG)训练的MAD策略能泛化到未见场景,性能与标准神经网络策略相当,且设计即保证闭环稳定性。

原文摘要 · Abstract (English)

We introduce magnitude and direction (MAD) policies, a policy parameterization for reinforcement learning (RL) that preserves Lp closed-loop stability for nonlinear dynamical systems. Despite their completeness in describing all stabilizing controllers, methods based on nonlinear Youla and system-level synthesis are significantly impacted by the difficulty of parametrizing Lp-stable operators. In contrast, MAD policies introduce explicit feedback on state-dependent features - a key element behind the success of reinforcement learning pipelines - without jeopardizing closed-loop stability. This is achieved by letting the magnitude of the control input be described by a disturbance-feedback Lp-stable operator, while selecting its direction based on state-dependent features through a universal function approximator. We further characterize the robust stability properties of MAD policies under model mismatch. Unlike existing disturbance-feedback policy parametrizations, MAD policies introduce state-feedback components compatible with model-free RL pipelines, ensuring closed-loop stability with no model information beyond assuming open-loop stability. Numerical experiments show that MAD policies trained with deep deterministic policy gradient (DDPG) methods generalize to unseen scenarios - matching the performance of standard neural network policies while guaranteeing closed-loop stability by design.

强化学习稳定性控制无模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。