用强化学习让机械臂在延迟约束下更省力、更安全地运行。
Reinforcement Learning-Based Neuroadaptive Control of Robotic Manipulators under Deferred Constraints
- 通过平滑约束机制,控制力度随接近约束逐渐增强。
- 无需精确建模,实时自适应调节,支持初始状态违规时安全运行。
- 适合高精度工业机械臂控制,尤其复杂动态环境应用。
本文提出一种基于强化学习的神经自适应控制框架,用于在延迟约束条件下运行的机器人机械臂。该方法改进了传统的障碍李雅普诺夫函数,引入平滑的约束执行机制,具有两大优势:(i) 在无约束区域最小化控制努力,接近约束时逐步增加,提升能效;(ii) 通过预设时间转移函数实现约束的渐进激活,允许初始状态违反约束时仍能安全操作。为应对系统不确定性并提升适应性,采用演员-评论家强化学习框架:评论家网络估计价值函数,演员网络实时学习最优控制策略,实现无需显式系统建模的自适应约束处理。基于李雅普诺夫的稳定性分析保证了闭环信号的有界性。通过数值仿真验证了所提方法的有效性。
原文摘要 · Abstract (English)
This paper presents a reinforcement learning-based neuroadaptive control framework for robotic manipulators operating under deferred constraints. The proposed approach improves traditional barrier Lyapunov functions by introducing a smooth constraint enforcement mechanism that offers two key advantages: (i) it minimizes control effort in unconstrained regions and progressively increases it near constraints, improving energy efficiency, and (ii) it enables gradual constraint activation through a prescribed-time shifting function, allowing safe operation even when initial conditions violate constraints. To address system uncertainties and improve adaptability, an actor-critic reinforcement learning framework is employed. The critic network estimates the value function, while the actor network learns an optimal control policy in real time, enabling adaptive constraint handling without requiring explicit system modeling. Lyapunov-based stability analysis guarantees the boundedness of all closed-loop signals. The effectiveness of the proposed method is validated through numerical simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。