研究KAN网络在持续学习中的遗忘问题,发现其在高维任务中仍易遗忘。
Catastrophic Forgetting in Kolmogorov-Arnold Networks
- 通过激活支持重叠与数据内在维度理论分析遗忘机制
- 实验显示低维任务保留好,高维任务如图像分类仍严重遗忘
- 提出KAN-LoRA适配器,适用于语言模型高效知识编辑
持续学习中的灾难性遗忘是长期挑战,模型在学习新任务时会丢失旧任务知识。尽管多层感知机已有多种缓解策略,但近年来提出的柯尔莫戈罗夫-阿诺德网络(KANs)因其基于局部样条激活,被认为具有内在抗遗忘能力。然而,KANs在持续学习中的实际表现尚不清晰,其局限性也未充分理解。为此,本文对KANs中的灾难性遗忘进行了全面研究,建立了将遗忘与激活支持重叠及内在数据维度关联的理论框架。通过在合成数据和视觉任务上的系统实验,测量了不同模型配置和数据复杂度下的遗忘动态。进一步提出KAN-LoRA,一种用于语言模型参数高效持续微调的新适配器设计,并在知识编辑任务中评估其有效性。结果表明,虽然KANs在低维算法场景中表现出良好记忆保持,但在高维领域如图像分类和语言建模中仍面临显著遗忘。这些发现深化了对KANs优劣的理解,为持续学习系统设计提供了实用洞见。
原文摘要 · Abstract (English)
Catastrophic forgetting is a longstanding challenge in continual learning, where models lose knowledge from earlier tasks when learning new ones. While various mitigation strategies have been proposed for Multi-Layer Perceptrons (MLPs), recent architectural advances like Kolmogorov-Arnold Networks (KANs) have been suggested to offer intrinsic resistance to forgetting by leveraging localized spline-based activations. However, the practical behavior of KANs under continual learning remains unclear, and their limitations are not well understood. To address this, we present a comprehensive study of catastrophic forgetting in KANs and develop a theoretical framework that links forgetting to activation support overlap and intrinsic data dimension. We validate these analyses through systematic experiments on synthetic and vision tasks, measuring forgetting dynamics under varying model configurations and data complexity. Further, we introduce KAN-LoRA, a novel adapter design for parameter-efficient continual fine-tuning of language models, and evaluate its effectiveness in knowledge editing tasks. Our findings reveal that while KANs exhibit promising retention in low-dimensional algorithmic settings, they remain vulnerable to forgetting in high-dimensional domains such as image classification and language modeling. These results advance the understanding of KANs' strengths and limitations, offering practical insights for continual learning system design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。