提出新方法让无人艇在洋流中快速适应,控制更平滑、更省电。
Knowledge-Distilled End-to-End Reinforcement Learning for Smooth 6-DOF Thrust Control and Rapid Adaptation to Ocean Currents in Remotely Operated Vehicles

- 两阶段知识蒸馏框架学习最优控制策略
- 稳态误差等五项指标降至传统方法的15%~43%
- 适合需要快速响应和节能的水下机器人应用
随着计算能力提升,端到端强化学习在遥控无人艇控制中快速发展。然而现有方法在洋流干扰下仍难以实现最优控制,缺乏统一框架同时兼顾低稳态跟踪误差、快速暂态响应、节能运行和控制力输出平滑性。为此,本文提出推力平滑快速洋流自适应近端策略优化(TSRCA-PPO)方法,通过双阶段知识蒸馏框架学习近优策略。核心创新在于奖励函数设计与特权多编码器架构。消融实验证明各模块有效性。仿真结果表明,所提方法在所有评估指标上均优于传统级联P-PID控制器:稳态位置误差、稳态姿态误差、调节时间、能耗指数和推力平滑指数分别降至对应P-PID值的42.7%、76.5%、10.6%、93.5%和15.9%。
原文摘要 · Abstract (English)
With the continuous improvement of computational capabilities, end-to-end reinforcement learning has been rapidly developed for remotely operated vehicles control. Nevertheless, existing end-to-end reinforcement-learningbased methods still face challenges in achieving optimal control under oceancurrent disturbances. In particular, there remains a lack of a unified control framework that can simultaneously achieve low steady-state tracking error, rapid transient response, energy-efficient operation, and smooth controlforce outputs under disturbances. To address the issue, this paper proposes the thrust smoothness rapid current adaptation proximal policy optimization (TSRCA-PPO) method which learns a near-optimal strategy by a twostage distillation learning framework. The core innovations of this work lie in the reward-function design and the privileged multi-encoder architecture. Ablation studies validate the effectiveness of each module. Simulation results demonstrate that the proposed TSRCA-PPO method consistently outperforms the conventional cascaded P-PID controller across all evaluation metrics. Specifically, TSRCA-PPO reduces the steady-state position error, steady-state attitude error, settling time, energy index, and thrustsmoothness index to 42.7%, 76.5%, 10.6%, 93.5%, and 15.9% of the corresponding P-PID values, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。