用分层强化学习实现异构机器人群体高精度编队控制
High-Precision Formation Control for Heterogeneous Multi-Robot Systems via Hierarchical Hybrid Physics-Informed Deep Reinforcement Learning

- 分层设计:上层用SAC算法控制领航车,下层融合物理模型与自适应强化学习
- 实验成功率100%,在扰动下仍保持稳定编队
- 适合需要高精度协同控制的机器人系统研发人员
现有经典控制方法通常依赖精确模型,难以应对模型不确定性和外部干扰;而端到端强化学习方法则存在样本效率低、收敛性差的问题。为此,本文提出一种分层混合物理信息深度强化学习(HHy-PIDRL)框架,旨在实现异构多机器人系统(HMRS)的高精度、高响应编队控制。该框架包含两层:首先,上层基于Soft Actor-Critic(SAC)算法设计面向阿克曼转向领航车的自主导航策略网络;其次,下层集成高保真物理前馈控制器、经典比例-微分(PD)控制器和自适应强化学习残差控制器,提出一种有效的混合模型与强化学习(HM-DRL)编队控制策略网络;第三,为全向跟随者设计独特的分层奖励函数,有效引导智能体学习精细化、稳定的控制策略。实验结果表明,上层自主导航策略网络与基于HM-DRL的编队控制策略网络的成功率均达到100%。同时,通过消融实验验证了所提方法的有效性与可靠性。
原文摘要 · Abstract (English)
Existing classical control methods commonly require precise models and struggle to cope with model uncertainties and external disturbances, while end-to-end reinforcement learning (RL) approaches suffer from low sample efficiency and poor convergence. To overcome these challenges, this paper proposes a hierarchical hybrid physics-informed deep reinforcement learning (HHy-PIDRL) framework, aiming to realize high-precision, highly responsive formation control for heterogeneous multi-robot systems (HMRSs). The proposed framework contains two layers. Specifically, first, the upper layer designs an autonomous navigation policy network for Ackermann-steering leader based on the Soft Actor-Critic (SAC) deep reinforcement learning (DRL) algorithm. Second, the lower module integrates a high-fidelity physical feed-forward controller, a classical proportional-derivative (PD) controller, and an adaptive DRL residual controller to propose an effective hybrid model and DRL (HM-DRL)-based formation control policy network. Third, a unique hierarchical reward function is designed for training Omnidirectional followers, which effectively guides agents toward a refined, stable control policy. Experimental results demonstrate that, the success rate of both the upper-layer autonomous navigation policy network and the HM-DRL based formation control policy networks reach 100%. Meanwhile, ablation experiments are conducted to verify the validity and credibility of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。