用强化学习协调无人船队,高效搜寻并清理水体塑料垃圾
Optimizing Plastic Waste Collection in Water Bodies Using Heterogeneous Autonomous Surface Vehicles with Deep Reinforcement Learning
- 分侦察与清理两队,通过深度强化学习协同决策
- 在复杂水域中比传统方法多清理30%以上垃圾
- 适合环境监测与海洋垃圾治理场景应用
本文提出一种无模型的深度强化学习框架,用于异构无人水面艇舰队的智能路径规划,以定位并收集水体中的塑料垃圾。系统由侦察队和清理队组成,通过深度强化学习实现两队协作,使侦察队持续更新污染分布模型,清理队据此高效捕捞垃圾。该策略通过定制奖励函数提升舰队整体效率。在高凸性区域与狭窄通道两种不同场景下,对比多种先进启发式算法,结果表明:基于深度强化学习的方法具备更强适应性,在复杂布局中尤其显著;采用贪婪动作训练可进一步提升性能。
原文摘要 · Abstract (English)
This paper presents a model-free deep reinforcement learning framework for informative path planning with heterogeneous fleets of autonomous surface vehicles to locate and collect plastic waste. The system employs two teams of vehicles: scouts and cleaners. Coordination between these teams is achieved through a deep reinforcement approach, allowing agents to learn strategies to maximize cleaning efficiency. The primary objective is for the scout team to provide an up-to-date contamination model, while the cleaner team collects as much waste as possible following this model. This strategy leads to heterogeneous teams that optimize fleet efficiency through inter-team cooperation supported by a tailored reward function. Different trainings of the proposed algorithm are compared with other state-of-the-art heuristics in two distinct scenarios, one with high convexity and another with narrow corridors and challenging access. According to the obtained results, it is demonstrated that deep reinforcement learning based algorithms outperform other benchmark heuristics, exhibiting superior adaptability. In addition, training with greedy actions further enhances performance, particularly in scenarios with intricate layouts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。