用深度强化学习实现6自由度水下机器人节能精准控制
Toward 6-DOF Autonomous Underwater Vehicle Energy-Aware Position Control based on Deep Reinforcement Learning: Preliminary Results
- 基于TQC算法直接控制推进器,无需手动调参
- 节能版方法平均省电30%,精度略低于传统控制器
- 适合需要长时续航的深海探测任务
自主水下航行器(AUV)在勘探、测绘和检查未知水下区域中至关重要,其机动性与能效是延长作业时间的关键。六自由度(6-DOF)全向平台因其高灵活性成为理想选择。尽管比例-积分-微分(PID)和模型预测控制(MPC)广泛应用,但需精确系统模型,对载荷或配置变化敏感,且调参耗时。现有基于深度强化学习(DRL)的方法多限于低自由度场景。本文提出一种新型基于TQC算法的DRL方法,用于控制全向6-DOF AUV,无需先验推进器配置知识,可直接输出推进器指令。同时将功耗纳入奖励函数。仿真结果显示,TQC高性能方法在抵达目标点时表现优于经调优的PID控制器;而TQC节能方法虽性能稍逊,但平均功耗降低30%。
原文摘要 · Abstract (English)
The use of autonomous underwater vehicles (AUVs) for surveying, mapping, and inspecting unexplored underwater areas plays a crucial role, where maneuverability and power efficiency are key factors for extending the use of these platforms, making six degrees of freedom (6-DOF) holonomic platforms essential tools. Although Proportional-Integral-Derivative (PID) and Model Predictive Control controllers are widely used in these applications, they often require accurate system knowledge, struggle with repeatability when facing payload or configuration changes, and can be time-consuming to fine-tune. While more advanced methods based on Deep Reinforcement Learning (DRL) have been proposed, they are typically limited to operating in fewer degrees of freedom. This paper proposes a novel DRL-based approach for controlling holonomic 6-DOF AUVs using the Truncated Quantile Critics (TQC) algorithm, which does not require manual tuning and directly feeds commands to the thrusters without prior knowledge of their configuration. Furthermore, it incorporates power consumption directly into the reward function. Simulation results show that the TQC High-Performance method achieves better performance to a fine-tuned PID controller when reaching a goal point, while the TQC Energy-Aware method demonstrates slightly lower performance but consumes 30% less power on average.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。