用大模型增强强化学习,让无人潜航器在恶劣海况下更智能地追踪目标
Never too Prim to Swim: An LLM-Enhanced RL-based Adaptive S-Surface Controller for AUVs under Extreme Sea Conditions
- 大模型协同优化强化学习的参数与奖励函数,实现多目标平衡
- 在复杂海况下完成目标追踪和数据采集任务,性能优于传统控制器
- 适合需要高适应性的水下自主系统研发人员参考
自主水下航行器(AUV)在海洋研究中备受关注,因其自由度间存在强耦合且易受不可预测干扰。本文提出一种基于大语言模型(LLM)增强的强化学习(RL)自适应S面控制器。通过引入多模态和结构化显式任务反馈,LLM实现控制器参数与奖励函数的联合优化,提升任务导向性能与适应性。所提控制器中,RL策略负责高层任务决策,输出面向任务的高级指令,由S面控制器转化为控制信号,有效抵消非线性效应及极端海况下的外部扰动。在含复杂地形、波浪与洋流的极端海况下,该控制器在水下目标追踪与数据采集等高层任务中表现出卓越性能与适应性,显著优于传统PID与SMC控制器。
原文摘要 · Abstract (English)
The adaptivity and maneuvering capabilities of Autonomous Underwater Vehicles (AUVs) have drawn significant attention in oceanic research, due to the unpredictable disturbances and strong coupling among the AUV's degrees of freedom. In this paper, we developed large language model (LLM)-enhanced reinforcement learning (RL)-based adaptive S-surface controller for AUVs. Specifically, LLMs are introduced for the joint optimization of controller parameters and reward functions in RL training. Using multi-modal and structured explicit task feedback, LLMs enable joint adjustments, balance multiple objectives, and enhance task-oriented performance and adaptability. In the proposed controller, the RL policy focuses on upper-level tasks, outputting task-oriented high-level commands that the S-surface controller then converts into control signals, ensuring cancellation of nonlinear effects and unpredictable external disturbances in extreme sea conditions. Under extreme sea conditions involving complex terrain, waves, and currents, the proposed controller demonstrates superior performance and adaptability in high-level tasks such as underwater target tracking and data collection, outperforming traditional PID and SMC controllers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。