在持续学习中,强制特征空间各向同性反而会降低模型性能。
Degradation of Feature Space in Continual Learning
- 通过对比学习在流式数据上训练,测试特征空间各向同性的影响。
- 在CIFAR-10和CIFAR-100上,各向同性正则化导致准确率下降。
- 揭示了集中式与持续学习中特征几何的本质差异,适合研究持续学习者。
集中式训练是深度学习的标准范式,使模型能在单一位置从统一数据集学习,自然产生各向同性的特征分布,以支持结构良好且可泛化的表示。相比之下,持续学习处理的是流式且非平稳的数据,采用增量方式训练,固有地面临塑性-稳定性困境。在此类设置下,学习动态倾向于产生越来越各向异性的特征空间。这引发一个根本问题:是否应强制各向同性以更好地平衡稳定性和塑性,从而缓解灾难性遗忘?本文研究在持续学习中促进特征空间各向同性能否提升表示质量。通过在CIFAR-10和CIFAR-100数据集上使用对比持续学习技术的实验,我们发现各向同性正则化不仅未能提升,反而会损害模型准确率。结果突显了集中式与持续学习在特征几何上的本质差异,表明各向同性虽在集中式设置中有益,但可能不适合作为非平稳学习场景的先验假设。
原文摘要 · Abstract (English)
Centralized training is the standard paradigm in deep learning, enabling models to learn from a unified dataset in a single location. In such setup, isotropic feature distributions naturally arise as a mean to support well-structured and generalizable representations. In contrast, continual learning operates on streaming and non-stationary data, and trains models incrementally, inherently facing the well-known plasticity-stability dilemma. In such settings, learning dynamics tends to yield increasingly anisotropic feature space. This arises a fundamental question: should isotropy be enforced to achieve a better balance between stability and plasticity, and thereby mitigate catastrophic forgetting? In this paper, we investigate whether promoting feature-space isotropy can enhance representation quality in continual learning. Through experiments using contrastive continual learning techniques on CIFAR-10 and CIFAR-100 data, we find that isotropic regularization fails to improve, and can in fact degrade, model accuracy in continual settings. Our results highlight essential differences in feature geometry between centralized and continual learning, suggesting that isotropy, while beneficial in centralized setups, may not constitute an appropriate inductive bias for non-stationary learning scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。