提出通用遗忘理论,揭示学习算法如何因预测不一致而丢失知识。
Forgetting is Everywhere
- 将遗忘定义为预测分布的自我不一致性导致的信息损失
- 证明贝叶斯推断可实现无遗忘的适应,且生成模型自训练时必然遗忘
- 在分类、回归、生成建模和强化学习中均验证遗忘普遍存在
开发通用学习算法的核心挑战在于其在适应新数据时会遗忘旧知识。尽管数十年研究,仍未形成统一的遗忘定义来揭示学习动态。本文提出一种算法与任务无关的理论,将遗忘视为学习者预测分布缺乏自洽性,表现为预测信息的流失。该理论自然导出衡量算法遗忘倾向的通用指标,证明精确贝叶斯推断可实现无遗忘的适应,并给出生成模型在自训练时必然遗忘的同义解释。通过涵盖分类、回归、生成建模和强化学习的全面实验验证,结果表明遗忘存在于所有深度学习场景,显著影响学习效率。这些发现为理解并提升通用学习算法的信息保留能力提供了原则性基础。
原文摘要 · Abstract (English)
A fundamental challenge in developing general learning algorithms is their tendency to forget past knowledge as they adapt to new data. Addressing this problem requires a principled understanding of forgetting. Yet, despite decades of study, no unified definition has emerged that offers insight into the underlying dynamics of learning. We propose an algorithm- and task-agnostic theory that characterises forgetting as a lack of self-consistency in a learner's predictive distribution, manifesting as a loss of predictive information. Our theory naturally yields a general measure of an algorithm's propensity to forget, proves that exact Bayesian inference allows for adaptation without forgetting, and provides a tautological explanation for why generative models forget when trained on their own synthetic outputs. To validate these claims, we design a comprehensive set of experiments that span classification, regression, generative modelling, and reinforcement learning. We demonstrate that forgetting is present across all deep learning settings and plays a significant role in determining learning efficiency. Together, these results establish a principled understanding of forgetting and lay the foundation for analysing and improving the information retention capabilities of general learning algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。