用CFD训练的代理模型提升水下机器人控制效率与鲁棒性
CORAL-AUV: CFD Oriented Reinforcement Learning for Autonomous Underwater Vehicles

- 用CFD数据训练代理拖曳模型,嵌入强化学习加速推理
- 能耗降31%,路径追踪快11%,误差减19%,零样本迁移成功
- 适合需要高精度物理建模的水下自主系统研发者
自主水下航行器(AUV)的精细控制与定位对采样、维护和勘测至关重要。传统控制方法依赖人工调参,且对构型或环境变化不鲁棒。强化学习(RL)通过领域随机化(DR)可快速开发控制器,但受限于仿真对真实物理的建模能力,尤其在拖曳力建模上存在显著的仿真到现实差距。计算流体动力学(CFD)虽能提供高保真拖曳模型,但因计算开销大,难以融入强化学习框架。本文首次提出训练特定车辆的拖曳力代理模型(SDM),基于CFD数据构建,可在RL中实现快速推理。我们成功部署了首个基于SDM的零样本6自由度AUV强化学习策略。相比使用简化物理的控制器,该策略在跨航点移动时能耗降低31%,速度提升11%,误差减少19%。其零样本迁移性能更优,且对奖励函数设计变化更具鲁棒性。在参数扰动任务中,仅有基于CFD的策略成功完成迁移。所有策略在受控水池和野外环境中均进行了验证。
原文摘要 · Abstract (English)
Fine grain control and positioning of autonomous underwater vehicles (AUVs) is critical for sampling, maintenance, and survey applications. Traditional control methods for AUVs are labor intensive and are not robust to changes in the vehicle configuration or environmental conditions. Reinforcement learning (RL) promises rapid controller development while handling a range of deployment parameters via domain randomization (DR). However, DR is still limited by the capacity of the underlying simulation to model real physics. In particular, drag physics are difficult to model and are a large contributor to sim-to-real gaps. Meanwhile, computational fluid dynamics (CFD) provides high fidelity drag models but is challenging to leverage within reinforcement learning frameworks due to its computational overhead. Thus, in this paper we exploit the idea of training surrogate approximations of CFD models of a given vehicle, enabling fast inference within RL pipelines. We are the first to successfully deploy a zero-shot RL policy on a 6-DOF AUV in which policy training is performed on surrogate drag models (SDMs) trained on CFD data. We find 31% lower energy usage compared to a controller using simplified physics while traversing between waypoints 11% faster with 19% less error. Our SDM based RL controller better predicts zero-shot transfer and is more robust across reward shaping design choices. When using DR to complete a task with perturbed parameters, we find that the CFD policy is the only controller that successfully transfers. The policies are evaluated in a controlled tank environment and in the field providing extensive testing of the policies' capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。