arXiv:2512.12855cs.ROcs.SY2025-12被引 1

融合MPC与强化学习,实现非线性系统的安全自适应控制

MPC-Guided Safe Reinforcement Learning and Lipschitz-Based Filtering for Structured Nonlinear Systems

  • 用MPC设定安全控制边界,指导RL学习
  • 部署时采用Lipschitz滤波器,实时保证约束满足
  • 适用于有界扰动的非线性系统,适合工程控制场景

现代工程系统(如自动驾驶车辆、柔性机器人、智能航空航天平台)需要在不确定性、环境变化和实时约束下具备鲁棒性、自适应性和安全性。强化学习(RL)能为具有非线性动态特性的系统提供强大的数据驱动自适应能力,但缺乏探索过程中的动态约束满足机制。模型预测控制(MPC)虽能结构化处理约束并保证鲁棒性,但依赖精确模型且在线优化计算量大,难以实时应用。本文提出一种集成的MPC-RL框架,将MPC的稳定性和安全性保障与RL的自适应能力结合。训练阶段,MPC定义安全控制边界以指导RL学习,实现约束感知策略;部署阶段,使用基于Lipschitz连续性的轻量级安全滤波器,无需复杂在线优化即可确保约束满足。该方法在非线性气动弹性机翼系统上验证,表现出更强的扰动抑制能力、更低的执行器能耗以及在湍流下的鲁棒性能。该架构可推广至其他具有结构性非线性与有界扰动的领域,为工程中基于人工智能的可控系统提供可扩展的安全解决方案。

原文摘要 · Abstract (English)

Modern engineering systems, such as autonomous vehicles, flexible robotics, and intelligent aerospace platforms, require controllers that are robust to uncertainties, adaptive to environmental changes, and safety-aware under real-time constraints. RL offers powerful data-driven adaptability for systems with nonlinear dynamics that interact with uncertain environments. RL, however, lacks built-in mechanisms for dynamic constraint satisfaction during exploration. MPC offers structured constraint handling and robustness, but its reliance on accurate models and computationally demanding online optimization may pose significant challenges. This paper proposes an integrated MPC-RL framework that combines stability and safety guarantees of MPC with the adaptability of RL. During training, MPC defines safe control bounds that guide the RL component and that enable constraint-aware policy learning. At deployment, the learned policy operates in real time with a lightweight safety filter based on Lipschitz continuity to ensure constraint satisfaction without heavy online optimizations. The approach, which is validated on a nonlinear aeroelastic wing system, demonstrates improved disturbance rejection, reduced actuator effort, and robust performance under turbulence. The architecture generalizes to other domains with structured nonlinearities and bounded disturbances, offering a scalable solution for safe artificial-intelligence-driven control in engineering applications.

强化学习安全控制MPC非线性系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。