arXiv:2510.22420cs.ROcs.SY2025-10被引 1

提出分层强化学习框架,实现高维随机系统的稳定自适应控制。

A Novel Multi-Timescale Stability-Preserving Hierarchical Reinforcement Learning Controller Framework for Adaptive Control in High-Dimensional Dynamical Systems

  • 分层策略结合半马尔可夫决策过程,高低层分工处理多时标决策。
  • 在8维超混沌系统和5自由度机械臂上,误差指标最低达IAE:1.623。
  • 理论保障稳定性,适合机器人、自动驾驶等复杂系统应用。

控制高维随机系统(如机器人、自动驾驶、超混沌系统)面临维度灾难、缺乏时间抽象及难以保证随机稳定性等问题。本文提出多时标李雅普诺夫约束分层强化学习(MTLHRL)框架,将分层策略嵌入半马尔可夫决策过程(SMDP),由高层策略进行战略规划,低层策略执行反应式控制,有效应对复杂多时标决策并降低维度开销。通过神经李雅普诺夫函数结合拉格朗日松弛与多时标演员-评论家更新,严格保证系统在随机动态下的均方有界性或渐近稳定性。框架引入信任域约束与解耦优化,提升学习效率与可靠性。在8维超混沌系统与5自由度机械臂上的大量仿真表明,该方法显著优于基线模型,在稳定性与性能上表现优异:超混沌控制中积分绝对误差(IAE)降至3.912,机器人控制中达1.623,收敛更快且抗干扰能力更强。MTLHRL为复杂随机系统的鲁棒控制提供了理论严谨且实用的解决方案。

原文摘要 · Abstract (English)

Controlling high-dimensional stochastic systems, critical in robotics, autonomous vehicles, and hyperchaotic systems, faces the curse of dimensionality, lacks temporal abstraction, and often fails to ensure stochastic stability. To overcome these limitations, this study introduces the Multi-Timescale Lyapunov-Constrained Hierarchical Reinforcement Learning (MTLHRL) framework. MTLHRL integrates a hierarchical policy within a semi-Markov Decision Process (SMDP), featuring a high-level policy for strategic planning and a low-level policy for reactive control, which effectively manages complex, multi-timescale decision-making and reduces dimensionality overhead. Stability is rigorously enforced using a neural Lyapunov function optimized via Lagrangian relaxation and multi-timescale actor-critic updates, ensuring mean-square boundedness or asymptotic stability in the face of stochastic dynamics. The framework promotes efficient and reliable learning through trust-region constraints and decoupled optimization. Extensive simulations on an 8D hyperchaotic system and a 5-DOF robotic manipulator demonstrate MTLHRL's empirical superiority. It significantly outperforms baseline methods in both stability and performance, recording the lowest error indices (e.g., Integral Absolute Error (IAE): 3.912 in hyperchaotic control and IAE: 1.623 in robotics), achieving faster convergence, and exhibiting superior disturbance rejection. MTLHRL offers a theoretically grounded and practically viable solution for robust control of complex stochastic systems.

强化学习稳定控制分层决策机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。