arXiv:2412.03959cs.ROcs.SY2024-12中稿 · IEEE Transactions …被引 10

FISHER用两阶段学习让多水下无人机更智能地追踪目标

Is FISHER All You Need in The Multi-AUV Underwater Target Tracking Task?

  • 先模仿专家行为学策略,再用新算法优化决策
  • 在多种场景下追踪成功率超90%,且能适应新任务
  • 适合做水下协同追踪的强化学习研究者

多自主水下航行器(AUV)协同执行水下目标追踪任务具有重要意义,但传统控制方法难以满足多样化需求。为此,我们提出一种两阶段示范学习框架FISHER,突出强化学习在该任务中的适应性,同时解决其对环境交互依赖大、奖励函数设计难等问题。第一阶段采用模仿学习(IL)改进生成对抗式算法,构建基于多智能体判别器-演员-评论家的架构,并引入基于纳什均衡的多智能体IL优化目标,实现策略提升并生成离线数据集。第二阶段提出多智能体独立广义决策变压器,通过分析高质量样本的潜在表示来匹配未来状态,而非依赖奖励函数,从而获得更强泛化能力的策略。此外,我们设计了“仿真到仿真”示范生成流程,利用传统控制方法生成专家示范,便于领域迁移。大量仿真实验表明,FISHER在多种场景下具备高稳定性、多任务性能和强泛化能力,平均追踪成功率超过90%。

原文摘要 · Abstract (English)

It is significant to employ multiple autonomous underwater vehicles (AUVs) to execute the underwater target tracking task collaboratively. However, it's pretty challenging to meet various prerequisites utilizing traditional control methods. Therefore, we propose an effective two-stage learning from demonstrations training framework, FISHER, to highlight the adaptability of reinforcement learning (RL) methods in the multi-AUV underwater target tracking task, while addressing its limitations such as extensive requirements for environmental interactions and the challenges in designing reward functions. The first stage utilizes imitation learning (IL) to realize policy improvement and generate offline datasets. To be specific, we introduce multi-agent discriminator-actor-critic based on improvements of the generative adversarial IL algorithm and multi-agent IL optimization objective derived from the Nash equilibrium condition. Then in the second stage, we develop multi-agent independent generalized decision transformer, which analyzes the latent representation to match the future states of high-quality samples rather than reward function, attaining further enhanced policies capable of handling various scenarios. Besides, we propose a simulation to simulation demonstration generation procedure to facilitate the generation of expert demonstrations in underwater environments, which capitalizes on traditional control methods and can easily accomplish the domain transfer to obtain demonstrations. Extensive simulation experiments from multiple scenarios showcase that FISHER possesses strong stability, multi-task performance and capability of generalization.

多智能体模仿学习水下追踪强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。