arXiv:2410.05899cs.LG2024-10被引 9

受大脑突触机制启发,提升预训练模型持续学习能力。

Brain-inspired continual pre-trained learner via silent synaptic consolidation

  • 模拟成熟大脑突触可塑性,平衡记忆稳定与学习灵活性。
  • 在增量学习任务中显著优于传统方法,保持高泛化性能。
  • 适合研究神经可塑性或生物启发式AI的学者参考。

预训练模型虽具强大泛化能力,但在增量学习新任务时仍易产生灾难性遗忘。现有基于架构的方法面临两大挑战:1)将预训练网络与可训练子网络融合,难以在动态任务中协调学习灵活性与记忆稳定性;2)预训练网络与各类子网络间缺乏强关联,影响推理时相关信息的有效提取。本文提出Artsy,受成熟大脑中通过尖峰时间依赖可塑性实现的静息突触激活机制启发,以增强预训练模型的持续学习能力。Artsy包含两个核心组件:训练阶段,模仿成熟大脑动态,维持预训练网络中已有知识的记忆稳定性,同时促进任务特异性子网络的学习灵活性;推理阶段,利用人工静息突触与功能突触,通过突触巩固建立预训练网络前突触神经元与子网络后突触神经元间的精准连接,从而高效提取测试样本中的相关知识。全面实验表明,该模型在类别增量学习任务上显著优于传统方法,同时提升了架构类方法的生物学可解释性。此外,Artsy为模拟生物突触机制提供了新路径,有望推动对人工与生物系统中神经可塑性的理解。

原文摘要 · Abstract (English)

Pre-trained models have demonstrated impressive generalization capabilities, yet they remain vulnerable to catastrophic forgetting when incrementally trained on new tasks. Existing architecture-based strategies encounter two primary challenges: 1) Integrating a pre-trained network with a trainable sub-network complicates the delicate balance between learning plasticity and memory stability across evolving tasks during learning. 2) The absence of robust interconnections between pre-trained networks and various sub-networks limits the effective retrieval of pertinent information during inference. In this study, we introduce the Artsy, inspired by the activation mechanisms of silent synapses via spike-timing-dependent plasticity observed in mature brains, to enhance the continual learning capabilities of pre-trained models. The Artsy integrates two key components: During training, the Artsy mimics mature brain dynamics by maintaining memory stability for previously learned knowledge within the pre-trained network while simultaneously promoting learning plasticity in task-specific sub-networks. During inference, artificial silent and functional synapses are utilized to establish precise connections between the pre-synaptic neurons in the pre-trained network and the post-synaptic neurons in the sub-networks, facilitated through synaptic consolidation, thereby enabling effective extraction of relevant information from test samples. Comprehensive experimental evaluations reveal that our model significantly outperforms conventional methods on class-incremental learning tasks, while also providing enhanced biological interpretability for architecture-based approaches. Moreover, we propose that the Artsy offers a promising avenue for simulating biological synaptic mechanisms, potentially advancing our understanding of neural plasticity in both artificial and biological systems.

持续学习生物启发预训练模型突触机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。