模型遗忘并非知识消失,可通过接口密钥恢复旧任务性能。
Forgetting is Not Erasure: Recovering Latent Knowledge via Transport Keys

- 用接口密钥对齐新旧模型中间特征,实现跨阶段知识传递。
- 在分拆CIFAR-100上,密钥使任务A性能恢复至原始水平。
- 适合研究持续学习中知识保留与重用的学者参考。
灾难性遗忘常被视作表征丢失:模型在顺序训练后似乎丧失了早期任务的特征。我们挑战这一更强版本的观点。在受控的持续学习设置中,我们发现大部分看似遗忘的现象源于内部阶段间接口漂移,而非任务相关计算的永久消失。通过一种拼接评估协议,将更新后网络的早期计算与前序网络的晚期计算结合,可选地由紧凑的任务特定运输密钥中介。运输密钥被描述为系统级接口对齐算子,基于少量成对锚点激活估计,并通过模型拼接验证。在使用ResNet风格网络的分拆CIFAR-100上,运输密钥在顺序训练任务B后恢复了任务A的绝大部分原始性能。在紧凑视觉变压器上也观察到类似恢复模式。结果表明,持续学习可能需要更优的索引与重访问机制,而不仅仅是防止权重变化的方法。
原文摘要 · Abstract (English)
Catastrophic forgetting is often framed as a representational problem: after sequential training, a model appears to lose the features that supported performance on earlier tasks. We challenge the stronger form of this view. Across controlled continual-learning settings, we find that a significant portion of apparent forgetting can be attributed to interface drift between internal stages rather than permanent erasure of task-relevant computation. We study this phenomenon through a stitched evaluation protocol that combines early computation from a post-update network with late computation from its predecessor, optionally mediated by a compact, task-specific transport key. We describe transport keys at a systems level as compact interface-alignment operators estimated from a small set of paired anchor activations and evaluated through model stitching. On split CIFAR-100 with a ResNet-style network, transport keys recover most of the original Task A performance after sequential training on Task B. On a compact vision transformer, we observe a similar recovery pattern. These results suggest that continual learning may require better mechanisms for indexing and re-accessing latent computations, not only methods that prevent weight change.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。