arXiv:2510.26347cs.LGcs.AI2025-10

改进强化学习算法,让无人潜航器在污染源稀疏随机的水下环境中更高效探测。

Reinforcement Learning for Pollution Detection in a Randomized, Sparse and Nonstationary Environment with an Autonomous Underwater Vehicle

  • 改造蒙特卡洛算法,加入多目标学习和位置记忆防重复探索
  • 新方法在稀疏奖励环境下探测成功率远超传统Q-learning和穷举搜索
  • 适合开发复杂水域中自主搜寻污染源的智能机器人系统

强化学习(RL)算法通过学习最大化奖励的动作来优化问题求解,但在随机和非平稳环境中面临巨大挑战。即使先进的RL算法也常难以应对此类条件。在利用自主水下航行器(AUV)搜寻水下污染云时,环境具有奖励稀疏性,多数动作导致零奖励。本文通过重新审视并改进经典RL方法,以实现对稀疏、随机且非平稳环境的高效适应。我们系统研究了大量改进策略,包括分层算法调整、多目标学习,以及引入位置记忆作为外部输出过滤器,防止状态重复访问。实验表明,改进后的蒙特卡洛方法显著优于传统Q-learning和两种穷举搜索策略,证明其在复杂环境中的适用潜力。结果表明,强化学习可通过合理改造有效应用于随机、非平稳和奖励稀疏场景。

原文摘要 · Abstract (English)

Reinforcement learning (RL) algorithms are designed to optimize problem-solving by learning actions that maximize rewards, a task that becomes particularly challenging in random and nonstationary environments. Even advanced RL algorithms are often limited in their ability to solve problems in these conditions. In applications such as searching for underwater pollution clouds with autonomous underwater vehicles (AUVs), RL algorithms must navigate reward-sparse environments, where actions frequently result in a zero reward. This paper aims to address these challenges by revisiting and modifying classical RL approaches to efficiently operate in sparse, randomized, and nonstationary environments. We systematically study a large number of modifications, including hierarchical algorithm changes, multigoal learning, and the integration of a location memory as an external output filter to prevent state revisits. Our results demonstrate that a modified Monte Carlo-based approach significantly outperforms traditional Q-learning and two exhaustive search patterns, illustrating its potential in adapting RL to complex environments. These findings suggest that reinforcement learning approaches can be effectively adapted for use in random, nonstationary, and reward-sparse environments.

强化学习水下探测稀疏奖励自主航行器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。