让无人船在复杂海况下更稳更安全,通过抗干扰强化学习提升导航性能。
Perturbation-mitigated USV Navigation with Distributionally Robust Reinforcement Learning
- 引入分布鲁棒优化,动态应对不同环境下的传感器噪声变化。
- 实测成功率达94.8%,碰撞率降低12.28%,节能27.99%。
- 适合海上自主航行系统研发与高风险场景下的鲁棒决策研究。
无人水面艇(USV)在未知复杂海洋环境中运行时,其鲁棒性至关重要,尤其是异方差观测噪声对基于传感器的导航任务构成严峻挑战。近年来,分布强化学习(DistRL)在无先验环境信息的情况下展现出良好表现。然而,现有方法忽略噪声模式随环境变化的问题,影响价值函数学习与安全导航。为此,本文提出DRIQN,将分布鲁棒优化(DRO)与隐式分位数网络结合,优化自然环境下的最坏情况性能。通过重放缓冲区中的显式子群体建模,DRIQN融合异质噪声源并聚焦关键鲁棒性场景。基于风险敏感环境的实验表明,相比最优基线方法,DRIQN实现+13.51%成功率、-12.28%碰撞率、+35.46%时间节省和+27.99%能耗降低。
原文摘要 · Abstract (English)
The robustness of Unmanned Surface Vehicles (USV) is crucial when facing unknown and complex marine environments, especially when heteroscedastic observational noise poses significant challenges to sensor-based navigation tasks. Recently, Distributional Reinforcement Learning (DistRL) has shown promising results in some challenging autonomous navigation tasks without prior environmental information. However, these methods overlook situations where noise patterns vary across different environmental conditions, hindering safe navigation and disrupting the learning of value functions. To address the problem, we propose DRIQN to integrate Distributionally Robust Optimization (DRO) with implicit quantile networks to optimize worst-case performance under natural environmental conditions. Leveraging explicit subgroup modeling in the replay buffer, DRIQN incorporates heterogeneous noise sources and target robustness-critical scenarios. Experimental results based on the risk-sensitive environment demonstrate that DRIQN significantly outperforms state-of-the-art methods, achieving +13.51\% success rate, -12.28\% collision rate and +35.46\% for time saving, +27.99\% for energy saving, compared with the runner-up.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。