arXiv:2605.25170cs.LGcs.AI2026-05中稿 · the Reinforcement …

提出可自适应生长修剪冻结的神经网络,提升机器人嗅觉导航能力

Grow-Prune-Freeze Networks: Adaptive & Continual Learning Technique for Olfactory Navigation

论文配图:Grow-Prune-Freeze Networks: Adaptive & Continual Learning Technique for Olfactory Navigation
图 1 · 摘自论文原文
  • 通过生长、修剪、冻结早期层实现持续学习
  • 在湍流气味流中达到94%导航成功率
  • 适合需要动态适应的机器人任务研究

嗅觉训练数据分散且非标准化,限制了世界模型的构建。嗅觉导航是高度动态和非平稳的任务,适合实时持续学习。我们提出一种名为生长-修剪-冻结(GPF)的自适应框架,使智能体能通过调整策略网络的早期层来响应世界复杂性变化。基于非线性随机矩阵理论,我们将Pennington与Worth(2017)的工作从单隐藏层扩展到多层持续学习模型,并证明网络权重的特征值结构在逐层添加时得以保持。基于期望SARSA的GPF在湍流气味流导航任务中达到94%的成功率,该任务具有部分可观测性和非平稳性,代表了机器人自适应学习中的“大世界”挑战。实验还表明GPF可推广至Atari强化学习、图像分类及自回归语言模型等任务。所有代码与数据已开源,以促进嗅觉机器人领域的研究。

原文摘要 · Abstract (English)

Training data for olfaction is scattered through disparate, non-standardized datasets that limit the ability to build representative world models. Olfactory navigation is a highly dynamic and non-stationary task that benefits from real-time continual learning. We introduce an adaptive framework called Grow-Prune-Freeze (GPF) networks that enable an agent to continually learn through growing, pruning, and freezing early layers of its policy in response to world complexity. Grounding GPFs in non-linear random matrix theory, we show that the work of Pennington & Worth (2017) can be extended from single hidden layers to n-layer continual-learning models, and that eigenvalue composition of network weights is preserved as successive layers are added. We show that GPFs based on Expected SARSA achieve a 94% success rate on turbulent plume navigation - a partially observable, non-stationary task representative of the "big world" challenges that motivate adaptive learning in robotics - and provide supporting methodology for applying GPFs in other world models. Further experiments amount evidence that GPFs may generalize well to other machine learning tasks such as reinforcement learning in Atari, image classification, and autoregressive language models. We open source all code and data to encourage improvements on and more research in olfactory robotics.

持续学习嗅觉导航神经网络机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。