arXiv:2608.20970cs.AI2026-08被引 1

发现深度模型会像记忆一样调用特征,突破了传统组合思维。

Deep Learning Models Also Recall Features

  • 用线性投影解释为输入激活下的特征调用
  • 跨架构通用,非仅限于语言模型
  • 为可解释性研究提供新分析工具

机制可解释性研究近年关注大语言模型如何从权重中回忆事实。本文认为,这种事实回忆指向更普遍的现象:我称之为‘特征召回’的深层学习通用操作。核心观察是,线性投影可被理解为在输入激活加权下检索存储信息。本文定义特征召回,证明其适用于多种架构,并与主流的特征组合范式相区分。还探讨了如何在机制层面识别特征召回案例。该理论为哲学家理解深度学习提供了新概念工具,也为可解释性研究指明了实证方向。

原文摘要 · Abstract (English)

Recent work in mechanistic interpretability has studied how large language models recall facts stored in their weights. This paper argues that factual recall points to something broader: a general kind of operation in deep learning models, which I call feature recall. The core observation is that a linear projection can be read as retrieving stored information scaled by input activations. I define feature recall, show it applies across architectures, and contrast it with the established paradigm of feature combination. I also consider how cases of feature recall might be mechanistically identified. The account gives philosophers a new conceptual tool for understanding deep learning, and points to empirical directions for mechanistic interpretability research.

可解释性特征召回深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。