arXiv:2409.17300cs.LGcs.AI2024-09

尝试用平滑最小值方法提升神经网络持续学习能力,但效果不明显。

Neural Network Plasticity and Loss Sharpness

  • 用尖锐度正则化寻找平滑极小值,试图增强模型适应新任务能力
  • 实验显示该方法对降低持续学习中的性能退化无显著作用
  • 适合关注持续学习与模型泛化关系的研究者

近年来,持续学习作为一种预测设置,因面向复杂、非平稳目标而日益受到关注。此类任务要求神经网络具备可塑性,即适应新任务的能力。近期研究表明,在非平稳强化学习框架中,新任务上的可塑性损失与损失曲面的尖锐度密切相关。本文探索了尖锐度正则化技术在该场景下的应用,这类方法在标准预测任务中被证明能提升泛化性能,旨在缓解可塑性损失。结果表明,这些技术在减少可塑性损失方面并无显著效果。

原文摘要 · Abstract (English)

In recent years, continual learning, a prediction setting in which the problem environment may evolve over time, has become an increasingly popular research field due to the framework's gearing towards complex, non-stationary objectives. Learning such objectives requires plasticity, or the ability of a neural network to adapt its predictions to a different task. Recent findings indicate that plasticity loss on new tasks is highly related to loss landscape sharpness in non-stationary RL frameworks. We explore the usage of sharpness regularization techniques, which seek out smooth minima and have been touted for their generalization capabilities in vanilla prediction settings, in efforts to combat plasticity loss. Our findings indicate that such techniques have no significant effect on reducing plasticity loss.

持续学习神经网络损失曲面可塑性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。