arXiv:2608.15854cs.LGcs.CV2026-08

发现表征位移是遗忘的几何标志,提出新方法缓解持续学习中的知识丢失。

Geometry of Forgetting: Representation Flux in Continual Learning

论文配图:Geometry of Forgetting: Representation Flux in Continual Learning
图 1 · 摘自论文原文
  • 引入表征通量衡量样本级表征变化,揭示遗忘的几何本质
  • 实验显示高通量会先于性能下降出现,且与置信度降低相关
  • 新方法FlowLess-R稳定表征,兼容多种重放框架,提升准确率

灾难性遗忘仍是持续学习的核心挑战,神经网络在学习新任务时会丢失旧知识。现有方法多通过参数正则化或经验回放缓解遗忘,但表征空间动态机制仍不清晰。本文研究序列学习中潜在表征的演化,提出表征通量——一种衡量训练过程中样本级表征位移的几何指标。结果显示,表征通量在多个基准上与灾难性遗忘强相关,时间分析表明高通量可先于性能下降出现。表征偏移还与置信度下降相关,而互补几何特性提供额外的样本级遗忘信息。基于此,我们提出FlowLess-R,一种表征空间正则化方法,通过约束回放表征相对于存储参考的位移,同时允许持续学习。该方法架构无关,可通过表征匹配项融入回放类方法。在SplitMNIST、SplitFashionMNIST、SplitCIFAR10和SplitTinyImageNet上的实验表明,结合ER、DER++和ER-ACE时,最终平均准确率提升,遗忘减少。结果表明表征通量是遗忘的有力几何标记,稳定潜在表征是一种简单有效的缓解策略。

原文摘要 · Abstract (English)

Catastrophic forgetting remains a fundamental obstacle to continual learning, where neural networks lose previously acquired knowledge while learning new tasks. Existing methods primarily mitigate forgetting through parameter regularization or experience replay, while the representation-space dynamics associated with forgetting remain less understood. We investigate latent representation evolution during sequential learning and introduce representation flux, a geometric measure of sample-level representation displacement across training. We show that representation flux is strongly associated with catastrophic forgetting across multiple benchmarks, with temporal analyses indicating that elevated flux can precede subsequent performance degradation. Representation displacement is also associated with confidence degradation, while complementary geometric properties provide additional information about sample-level forgetting. Motivated by these observations, we propose FlowLess-R, a representation-space regularization method that constrains replay representations relative to stored references while allowing continued learning. FlowLess-R is architecture-agnostic and integrates into replay-based methods through a representation-matching term. Experiments on SplitMNIST, SplitFashionMNIST, SplitCIFAR10, and SplitTinyImageNet show improved final average accuracy and reduced forgetting with ER, DER++, and ER-ACE. Our results identify representation flux as an informative geometric marker of forgetting and show that stabilizing latent representations provides a simple strategy for mitigating catastrophic forgetting.

持续学习表征稳定几何分析遗忘缓解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。