arXiv:2604.21640cs.LGcs.AI2026-04

发现强化学习模型中仅1.5%权重负责区分水下导航任务,且多数连接输入上下文变量。

Task-specific Subnetwork Discovery in Reinforcement Learning for Autonomous Underwater Navigation

论文配图:Task-specific Subnetwork Discovery in Reinforcement Learning for Autonomous Underwater Navigation
图 1 · 摘自论文原文
  • 通过分析多任务强化学习网络内部结构,定位任务专属子网络。
  • 在相关任务中,仅约1.5%权重用于任务区分,85%连接输入上下文节点。
  • 结果有助于提升模型可解释性,适用于水下监测的持续学习与编辑。

自主水下航行器需在动态、不确定且感知受限条件下,以可解释方式适应执行多种任务,这是传统控制器难以应对的挑战。多任务强化学习通过共享表征实现跨任务高效适应,但其策略仍缺乏透明性,难以揭示内部决策机制,制约实际部署。本文在HoloOcean模拟器中分析预训练多任务强化学习网络的内部结构,识别并比较导航至不同物种的任务专属子网络。研究发现,在相关任务的上下文多任务设置中,仅约1.5%的权重用于任务区分,其中约85%连接输入层的上下文变量节点与下一隐藏层,凸显上下文信息的关键作用。该方法揭示了共享与专用网络组件,为基于上下文的多任务强化学习在水下长期监测中的高效模型编辑、迁移学习与持续学习提供支持。

原文摘要 · Abstract (English)

Autonomous underwater vehicles are required to perform multiple tasks adaptively and in an explainable manner under dynamic, uncertain conditions and limited sensing, challenges that classical controllers struggle to address. This demands robust, generalizable, and inherently interpretable control policies for reliable long-term monitoring. Reinforcement learning, particularly multi-task RL, overcomes these limitations by leveraging shared representations to enable efficient adaptation across tasks and environments. However, while such policies show promising results in simulation and controlled experiments, they yet remain opaque and offer limited insight into the agent's internal decision-making, creating gaps in transparency, trust, and safety that hinder real-world deployment. The internal policy structure and task-specific specialization remain poorly understood. To address these gaps, we analyze the internal structure of a pretrained multi-task reinforcement learning network in the HoloOcean simulator for underwater navigation by identifying and comparing task-specific subnetworks responsible for navigating toward different species. We find that in a contextual multi-task reinforcement learning setting with related tasks, the network uses only about 1.5% of its weights to differentiate between tasks. Of these, approximately 85% connect the context-variable nodes in the input layer to the next hidden layer, highlighting the importance of context variables in such settings. Our approach provides insights into shared and specialized network components, useful for efficient model editing, transfer learning, and continual learning for underwater monitoring through a contextual multi-task reinforcement learning method.

强化学习可解释性水下导航子网络发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。