arXiv:2605.06834cs.LG2026-05

提出新方法评估神经元重要性,解决深度网络持续学习时难更新的问题。

Attribution-Based Neuron Utility for Plasticity Restoration in Deep Networks

论文配图:Attribution-Based Neuron Utility for Plasticity Restoration in Deep Networks
图 1 · 摘自论文原文
  • 基于参考梯度差计算神经元功能代价,精准衡量重置价值
  • 在连续学习场景中显著提升模型可塑性,避免性能退化
  • 适合研究持续学习、模型可塑性与参数优化的学者

持续学习旨在同时保持新知识获取与旧知识保存能力。尽管灾难性遗忘广受关注,但深层网络在持续训练中还会因可塑性下降而难以更新,表现为神经元饱和、参数范数增长和有用曲率方向丧失。已有自适应重置策略通过重置低效参数来恢复训练能力,但其依赖的激活强度、贡献度或梯度活动等代理信号常与重置目标错位。本文提出梯度与参考值之差(GXD),一种基于参考梯度归因的理论驱动型实用度量,用于估计单元替换带来的首阶函数代价。实验表明,与重置功能代价对齐的度量能显著提升干预可靠性,在现有标准失效的场景下表现更优。GXD将自适应重置重构为干预成本估计问题,为构建更鲁棒的持续学习系统提供可行路径。

原文摘要 · Abstract (English)

Continual learning research attempts to conserve two fundamental capabilities: new knowledge acquisition and the preservation of previously acquired knowledge. While knowledge in this case can be measured through performance over an implicit or explicit task space, model plasticity generally concerns adaptability as data distributions evolve. Though much of the literature has focused on catastrophic forgetting, deep networks can also suffer from loss of plasticity, becoming progressively harder to update under continued training. Recent research has identified multiple mechanisms underlying this phenomenon, including neuron saturation, parameter norm growth, and loss of useful curvature directions. Adaptive reset-based interventions, which selectively reinitialize low-utility network parameters, have emerged as practical solutions to restore trainability. Existing utility measures used to guide resets, such as activation magnitude, contribution utility, or gradient-based activity, rely on proxy signals that can become misaligned with the intervention they are meant to guide. In this paper, we introduce gradient times difference from reference (GXD), a theoretically motivated utility measure based on reference-based gradient attribution that estimates the first-order functional cost of replacing a unit. Our results show that utility measures aligned with the functional cost of the reset can make interventions more reliable in settings where existing reset criteria degrade. GXD reframes adaptive resetting as an intervention cost estimation problem, providing a practical path toward more robust continual learning systems.

持续学习神经元评估可塑性重置机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。