arXiv:2606.14997cs.AIcs.LG2026-06中稿 · ICML

给深度学习模型找到类脑记忆痕迹,可精准删改知识。

AI Engram: In Search of Memory Traces in Artificial Intelligence

论文配图:AI Engram: In Search of Memory Traces in Artificial Intelligence
图 1 · 摘自论文原文
  • 用几何方法定义并提取神经网络中的记忆单元
  • 实验证明可线性操作任意记忆,无需重新训练
  • 适合研究模型可解释性与知识控制的学者

记忆形成是智能的基础,但深度神经网络是否保留可识别的记忆痕迹,类似生物记忆单元,仍是开放问题。本文提出一种几何框架,将神经科学中的特异性、可重激活、充分性与必要性标准形式化为约束逆问题。推导出闭式估计器,能从全局纠缠的参数中分离出单个记忆痕迹,并发现该解对应于参数流形上的自然梯度更新。由此定义的AI记忆痕迹可实现对已学知识的外科式操控:通过线性运算即可组合或擦除任意记忆子集,无需迭代优化。从简单MLP到大语言模型的实验表明,该方法具有因果有效性与显著可扩展性。结果连接了生物记忆理论与人工表征学习,揭示深度网络如何在分布式存储中同时实现功能特异性。

原文摘要 · Abstract (English)

Memory formation is fundamental to intelligence, yet whether deep neural networks preserve identifiable memory traces analogous to biological memory units remains an open question. This work introduces a geometric framework to identify such "AI engrams" by formalizing the neuroscientific criteria of specificity, reactivation, sufficiency, and necessity into a constrained inverse problem. We derive a closed-form estimator that isolates individual memory traces from globally entangled parameters, and show that this biologically-derived solution corresponds to a natural gradient update on the parameter manifold. AI engrams enable surgical manipulation of learned knowledge: any subset of memories can be composed or erased through linear arithmetic, without iterative optimization. Experiments ranging from simple MLPs to LLMs demonstrate the causal validity and substantial scalability of AI engrams. Together, these results bridge theories of biological memory and artificial representation learning and offer geometric insight into how deep networks simultaneously support functional specificity within distributed storage.

记忆机制可解释性知识编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。