多智能体强化学习实现无目标状态估计的协同追击,实现实时稳定追踪。
Cooperative Bearing-Only Target Pursuit via Multiagent Reinforcement Learning: Design and Experiment
- 基于纯方位信息设计统一滤波器,提升目标估计稳定性。
- 提出新型多智能体强化学习框架,支持异构车辆在复杂环境追击。
- 通过可调控制增益与谱归一化算法,实现仿真到真实机器人的零样本迁移。
本文针对未知目标的多机器人追击问题,同时解决目标状态估计与追击控制。在状态估计方面,仅利用视觉传感器提供的方位信息,针对方位测量非线性带来的不稳定及双角度表示的奇异性,提出统一的纯方位信息滤波器,融合多维方位测量,形式简洁,增强稳定性并提高对视场受限导致目标丢失的鲁棒性。在追击控制方面,针对异构性与视场受限等挑战,传统方法如微分博弈或Voronoi划分常显不足,本文提出新型多智能体强化学习(MARL)框架,使多个异构车辆能协同搜索、定位并追踪目标。为弥合仿真到现实的差距,引入两项关键技术:训练中加入可调低层控制增益以模拟真实自主地面车辆(AGVs)动力学,以及采用谱归一化强化学习算法提升策略平滑性与鲁棒性。最终成功实现MARL控制器在实际AGVs上的零样本迁移,验证了方法的有效性与实用性。相关视频见 https://youtu.be/HO7FJyZiJ3E。
原文摘要 · Abstract (English)
This paper addresses the multi-robot pursuit problem for an unknown target, encompassing both target state estimation and pursuit control. First, in state estimation, we focus on using only bearing information, as it is readily available from vision sensors and effective for small, distant targets. Challenges such as instability due to the nonlinearity of bearing measurements and singularities in the two-angle representation are addressed through a proposed uniform bearing-only information filter. This filter integrates multiple 3D bearing measurements, provides a concise formulation, and enhances stability and resilience to target loss caused by limited field of view (FoV). Second, in target pursuit control within complex environments, where challenges such as heterogeneity and limited FoV arise, conventional methods like differential games or Voronoi partitioning often prove inadequate. To address these limitations, we propose a novel multiagent reinforcement learning (MARL) framework, enabling multiple heterogeneous vehicles to search, localize, and follow a target while effectively handling those challenges. Third, to bridge the sim-to-real gap, we propose two key techniques: incorporating adjustable low-level control gains in training to replicate the dynamics of real-world autonomous ground vehicles (AGVs), and proposing spectral-normalized RL algorithms to enhance policy smoothness and robustness. Finally, we demonstrate the successful zero-shot transfer of the MARL controllers to AGVs, validating the effectiveness and practical feasibility of our approach. The accompanying video is available at https://youtu.be/HO7FJyZiJ3E.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。