实验证明顺序删记忆会导致结果偏差,线性记忆模型不成立。
Forgetting Is Not a Fix: Path Dependence in Sequential Engram Editing
- 在连续编辑中测试记忆删除效果,发现顺序影响结果
- 删减顺序不同导致71%的编辑效果差异,且知识会部分恢复
- 适合关注模型可解释性与持续学习的研究者
AI Engram(Kwon等,2026)将神经科学中的记忆痕迹标准形式化为权重空间的约束逆问题,并以闭式解求解:特定概念的记忆痕迹成为线性对象,可一次性提取并算术组合。附录F提出组合记忆状态假说:编辑后的模型位于‘交换性流形’上,无论学习顺序如何,概念A与B的整合都能达到一致平衡。但该假说仅基于单次或成对编辑(即单周期测试),无法反映疲劳累积的时序动态。本文在作者提供的参考实现上,使用报告最优编辑强度(TOFU alpha=0.6)进行预注册测试,覆盖三个模型来源(两家供应商,两种架构)。四个发现跨所有测试均复现:(1) 零样本组合与序列重校准编辑间差异达编辑幅度的61%-71%;(2) 删减顺序不可互换,影响随概念重叠度增加——某次测试中,两个巴黎地标删减顺序决定了第三个无关概念是否存活;(3) 所有存活概念的层输入协方差(方法自身充分统计量,视为应变计)随每次删减单调漂移;(4) 被擦除知识在后续无关删减下部分回归。因此,附录F的交换性流形假说被序列编辑场景证伪。原始论文的单次编辑结果仍有效。对于合规性删除而言,今日认证的擦除,无法保证模型下次编辑后的状态。
原文摘要 · Abstract (English)
AI Engram (Kwon et al., 2026) formalizes the four engram criteria of neuroscience as a constrained inverse problem in weight space and solves it closed-form: concept-specific memory traces become linear objects that can be extracted once and combined arithmetically. Appendix F states the Compositional Memory States Hypothesis: edited models live on "a commutative manifold where the integration of A and B reaches a consistent equilibrium regardless of the learning sequence." The evidence base is single and paired edits -- in materials terms, single-cycle tests, in which fatigue accumulation is structurally invisible. Whether the hypothesis holds under sequential load is exactly the "temporal dynamics" question the paper defers to future work. We run that test on the authors' own reference implementation, at their reported best edit strength (TOFU alpha=0.6, a choice favoring the linearity hypothesis), with pre-registered predictions, across three model charges (two vendors, two architecture families). Four findings replicate across all three: (1) zero-shot composition and sequential re-calibrated editing diverge by 61-71% of the edit magnitude; (2) cut order is not interchangeable, and the effect scales with concept overlap -- in one charge the order of cutting two Paris landmarks decides whether an uninvolved third concept survives; (3) the survivors' layer-input covariances -- the method's own sufficient statistics, read as strain gauges -- drift monotonically with every further cut, in every surviving concept, in every charge; (4) erased knowledge partially returns under subsequent unrelated cuts. Appendix F's commutative-manifold hypothesis is thereby falsified for sequential editing; the single-edit results of the original paper are untouched. For unlearning-as-compliance: erasure certified today does not certify the artifact after its next edit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。