arXiv:2603.00043cs.LGcs.AI2026-03被引 2

用有限数据实现强化学习控制的稳定保证,让算法更可靠。

Reinforcement Learning for Control with Probabilistic Stability Guarantee: A Finite-Sample Approach

  • 基于李雅普诺夫方法,用有限轨迹推导概率稳定性定理。
  • 数据越多越长,稳定概率越高,最终趋于确定性。
  • 提出L-REINFORCE算法,适合需要稳定性的控制场景。

本文提出一种新型强化学习控制方法,可在仅使用有限采样轨迹的情况下提供概率稳定性保障。通过引入李雅普诺夫方法,我们建立了基于有限数据的概率稳定性定理,确保均方稳定性。稳定概率随轨迹数量和长度增加而提升,随着数据量增大趋于确定性。此外,我们推导了用于稳定策略学习的策略梯度定理,并开发了扩展经典REINFORCE算法的L-REINFORCE算法。在Cartpole任务上的仿真结果表明,该算法在保证稳定性方面优于基线方法。本工作弥合了强化学习与控制理论之间的关键差距,实现了在模型无关框架下基于有限数据的稳定性分析与控制器设计。

原文摘要 · Abstract (English)

This paper presents a novel approach to reinforcement learning (RL) for control systems that provides probabilistic stability guarantees using finite data. Leveraging Lyapunov's method, we propose a probabilistic stability theorem that ensures mean square stability using only a finite number of sampled trajectories. The probability of stability increases with the number and length of trajectories, converging to certainty as data size grows. Additionally, we derive a policy gradient theorem for stabilizing policy learning and develop an RL algorithm, L-REINFORCE, that extends the classical REINFORCE algorithm to stabilization problems. The effectiveness of L-REINFORCE is demonstrated through simulations on a Cartpole task, where it outperforms the baseline in ensuring stability. This work bridges a critical gap between RL and control theory, enabling stability analysis and controller design in a model-free framework with finite data.

强化学习控制理论稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。