发现数据同时观察能显著提升模型持续学习能力,超越记忆保留本身。
Forgetting, plasticity, and co-observation: a third facet of continual learning

- 分离训练时引入数据共观察机制,解耦数据访问约束与稳定性、可塑性。
- 在监督与自监督任务中,联合训练始终优于分步训练,性能差距稳定存在。
- 记忆回放的成功不仅在于防止遗忘,更在于恢复数据共观察带来的优势。
高效持续学习仍是深度神经网络的核心挑战。尽管灾难性遗忘和可塑性丧失常被视为主要障碍,我们发现二者无法完全解释简单顺序训练与离线联合训练之间的性能差距。本文揭示数据共观察是影响持续学习表现的独立因素。通过将数据独立访问的限制与稳定性和可塑性解耦,我们系统研究了共同观察训练数据所带来的表征收益。实证表明,在通用的数据增量‘分块’场景下,无论监督还是自监督范式,联合训练与分步训练间均存在一致的性能差异,且在缓解遗忘、控制可塑性的前提下依然成立。结果表明,同时观察训练数据(共观察)能带来超越单纯知识保留的泛化优势,且该效应不依赖特定的持续分布变化。此外,我们以此视角重新审视主流持续学习机制:基于知识蒸馏的方法仅具有效果良好的知识保留能力;而记忆回放的实证成功远超遗忘缓解,其核心在于主动将数据共观察的优势重新引入学习过程。
原文摘要 · Abstract (English)
Efficient continual learning remains a fundamental challenge for deep neural networks. While catastrophic forgetting and loss of plasticity are widely considered the primary obstacles to overcome, we show that these two issues cannot fully explain the performance gap between naive sequential training and offline joint training. In this paper, we highlight data co-observation as a distinct factor influencing continual learning performance. By decoupling the constraints of separate data access from stability and plasticity, we systematically investigate the representational benefits gained by observing training data together. Empirically, we demonstrate a consistent performance difference between joint and separate training across both supervised and self-supervised paradigms in generic data-incremental "chunking" scenarios, whilst mitigating forgetting and controlling for plasticity. Our findings indicate that simultaneous observation of training data (co-observation) yields benefits to the learner's generalization that extend well beyond mere knowledge retention, and that this effect does not require a specific continual distribution shift. Furthermore, we contextualize prominent continual learning mechanisms through this lens: while distillation-based approaches act only as effective knowledge retention mechanisms, our results suggest that the empirical success of memory replay goes beyond the mitigation of forgetting, actively reintroducing the benefits of data co-observation into the learning process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。