突破注意力模型记忆容量的理论限制,实现任意上下文下的精准记忆。
Memorization in Attention-only Transformers
- 提出新证明方法,适用于任意上下文长度的注意力模型。
- 在单层注意力中实现更高效的精确记忆,且支持分布近似记忆。
- 实验验证理论边界更贴近真实模型能力,适合理论研究者参考。
近期研究探讨了多头注意力的记忆能力,但受限于不现实的上下文长度假设。本文提出一种针对基于语言的Transformer的新证明,将现有假说推广至任意上下文大小。该方法在单注意力层中实现了更有效的精确记忆,同时引入了分布近似记忆的概念。通过实验验证,我们表明所提界限更准确地反映了语言模型的真实记忆容量,并与先前工作提供了精确对比。
原文摘要 · Abstract (English)
Recent research has explored the memorization capacity of multi-head attention, but these findings are constrained by unrealistic limitations on the context size. We present a novel proof for language-based Transformers that extends the current hypothesis to any context size. Our approach improves upon the state-of-the-art by achieving more effective exact memorization with an attention layer, while also introducing the concept of approximate memorization of distributions. Through experimental validation, we demonstrate that our proposed bounds more accurately reflect the true memorization capacity of language models, and provide a precise comparison with prior work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。