arXiv:2502.05164cs.LGcond-mat.dis-nn2025-02ICML被引 11

单层Transformer可实现上下文去噪,揭示注意力与联想记忆的深层联系。

In-context denoising with one-layer transformers: connections between attention and associative memory retrieval

  • 用贝叶斯框架证明单层Transformer能最优解特定去噪问题。
  • 注意力层通过一步梯度更新优化上下文能量景观,优于直接检索。
  • 为上下文学习研究提供联想记忆新视角,适合模型机制探索者。

我们提出上下文去噪任务,将基于注意力的架构与密集联想记忆(DAM)网络(即现代霍普菲尔德网络)联系起来。通过贝叶斯框架,我们从理论上和实验上证明,某些受限的去噪问题即使仅用单层Transformer也能最优求解。我们表明,训练后的注意力层在处理每个去噪提示时,会基于上下文感知的DAM能量景观执行一次梯度下降更新:上下文标记作为关联记忆,查询标记作为初始状态。这一单步更新得到的解优于直接检索任一上下文标记或错误局部极小值,为DAM网络突破标准检索范式提供了具体例证。本工作强化了Ramsauer等人首次识别的联想记忆与注意力机制之间的联系,并展示了联想记忆模型在上下文学习研究中的相关性。

原文摘要 · Abstract (English)

We introduce in-context denoising, a task that refines the connection between attention-based architectures and dense associative memory (DAM) networks, also known as modern Hopfield networks. Using a Bayesian framework, we show theoretically and empirically that certain restricted denoising problems can be solved optimally even by a single-layer transformer. We demonstrate that a trained attention layer processes each denoising prompt by performing a single gradient descent update on a context-aware DAM energy landscape, where context tokens serve as associative memories and the query token acts as an initial state. This one-step update yields better solutions than exact retrieval of either a context token or a spurious local minimum, providing a concrete example of DAM networks extending beyond the standard retrieval paradigm. Overall, this work solidifies the link between associative memory and attention mechanisms first identified by Ramsauer et al., and demonstrates the relevance of associative memory models in the study of in-context learning.

注意力机制联想记忆上下文学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。