arXiv:2512.15735math.OCcs.AI2025-12被引 4

用强化学习+事件触发,实现不确定非线性系统的高效鲁棒控制

Deep Reinforcement Learning Optimization for Uncertain Nonlinear Systems via Event-Triggered Robust Adaptive Dynamic Programming

  • 结合强化学习与扰动观测器,实时估计状态和干扰
  • 事件触发机制使参数更新仅在偏差超标时进行,降低计算开销
  • 无需精确模型即可逼近最优控制,适合资源受限场景

本文提出一种统一控制架构,将强化学习驱动的控制器与扰动抑制的扩展状态观测器(ESO)结合,并引入事件触发机制(ETM)以减少不必要的计算。ESO 实时估计系统状态和总扰动,为有效补偿提供基础。为在无精确系统模型情况下获得近似最优行为,采用基于值迭代的自适应动态规划(ADP)方法进行策略逼近。事件触发机制确保学习模块的参数更新仅在状态偏差超过预设阈值时执行,从而避免过度学习活动,显著降低计算负荷。通过李雅普诺夫分析证明闭环系统的稳定性。数值实验进一步验证,该方法在保持强控制性能和扰动容忍能力的同时,相较于标准时间触发的 ADP 方案,大幅减少了采样与处理负担。

原文摘要 · Abstract (English)

This work proposes a unified control architecture that couples a Reinforcement Learning (RL)-driven controller with a disturbance-rejection Extended State Observer (ESO), complemented by an Event-Triggered Mechanism (ETM) to limit unnecessary computations. The ESO is utilized to estimate the system states and the lumped disturbance in real time, forming the foundation for effective disturbance compensation. To obtain near-optimal behavior without an accurate system description, a value-iteration-based Adaptive Dynamic Programming (ADP) method is adopted for policy approximation. The inclusion of the ETM ensures that parameter updates of the learning module are executed only when the state deviation surpasses a predefined bound, thereby preventing excessive learning activity and substantially reducing computational load. A Lyapunov-oriented analysis is used to characterize the stability properties of the resulting closed-loop system. Numerical experiments further confirm that the developed approach maintains strong control performance and disturbance tolerance, while achieving a significant reduction in sampling and processing effort compared with standard time-triggered ADP schemes.

强化学习鲁棒控制事件触发自适应动态规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。