揭示生成模型持续训练中遗忘的定量机制,解析为何旧知识会丢失或偏移。
A Quantitative Characterization of Forgetting in Post-Training
- 基于双模态混合模型,区分质量遗忘与旧成分漂移两种遗忘形式。
- 反向KL可避免质量遗忘,旧成分漂移随模式分离指数衰减。
- 揭示重放策略如何影响不同目标函数下的遗忘行为,适合研究持续学习者。
持续后训练广泛用于生成模型,但对其遗忘发生时机与原因仍缺乏系统理解。本文在Chen等人(2025)提出的双模态混合抽象框架下,形式化了两种遗忘:(i) 质量遗忘——旧模态权重降为零;(ii) 旧成分漂移——已正确旧成分在训练中发生偏移。对于等协方差高斯模态,证明正向KL目标在新分布数据上训练会使旧权重趋零,而反向KL目标收敛至真实目标(避免质量遗忘),且旧均值扰动仅由重叠门控误分配概率控制,该概率受巴塔查里亚系数调节,漂移随模态分离指数衰减,并具有局部良好条件几何与指数收敛性。进一步量化重放策略的作用:对正向KL,重放需改变训练分布以改变总体最优解;对反向KL,重放不改变总体目标,但通过有界重要性加权防止有限批次下的旧模态饥饿。最后,通过同一框架分析三种近期近策略后训练方法(SDFT、TTT-Discover、OAPL),推导出每种方法保持旧质量并呈现重叠控制漂移的显式条件。总体表明,遗忘可精确量化为发散方向、几何重叠、采样方式及训练中过往行为可见性之间的相互作用。
原文摘要 · Abstract (English)
Continual post-training of generative models is widely used, yet a principled understanding of when and why forgetting occurs remains limited. We develop theoretical results under a two-mode mixture abstraction (representing old and new tasks), proposed by Chen et al. (2025) (arXiv:2510.18874), and formalize forgetting in two forms: (i) mass forgetting, where the old mixture weight collapses to zero, and (ii) old-component drift, where an already-correct old component shifts during training. For equal-covariance Gaussian modes, we prove that forward-KL objectives trained on data from the new distribution drive the old weight to zero, while reverse-KL objectives converge to the true target (thereby avoiding mass forgetting) and perturb the old mean only through overlap-gated misassignment probabilities controlled by the Bhattacharyya coefficient, yielding drift that decays exponentially with mode separation and a locally well-conditioned geometry with exponential convergence. We further quantify how replay interacts with these objectives. For forward-KL, replay must modify the training distribution to change the population optimum; for reverse-KL, replay leaves the population objective unchanged but prevents finite-batch old-mode starvation through bounded importance weighting. Finally, we analyze three recently proposed near-on-policy post-training methods, SDFT (arxiv:2601.19897), TTT-Discover (arxiv:2601.16175), and OAPL (arxiv:2602.19362), via the same lens and derive explicit conditions under which each retains old mass and exhibits overlap-controlled drift. Overall, our results show that forgetting can by precisely quantified based on the interaction between divergence direction, geometric behavioral overlap, sampling regime, and the visibility of past behavior during training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。