arXiv:2508.08289cs.LGcs.AI2025-08被引 1

线性注意力机制可被解读为动物学习理论中的条件反射模型,解释其记忆行为。

Attention as Conditioning: What Classical Learning Theory Predicts About Linear Transformers

  • 将线性注意力映射到经典动物学习模型,揭示其内在机制
  • 模拟验证卡明阻断效应误差低于10^-7,与理论完全一致
  • 揭示注意力头提供容量而非冗余,适合研究模型内部机制

注意力常被视为关联记忆,但无法预测其行为。我们发现主流线性注意力家族的状态更新与百年动物学习理论中的命名模型完全对应:线性注意力实现海布相邻性,DeltaNet实现雷斯科拉-沃格纳误差修正,衰减变体如RetNet实现带刺激痕迹的相邻性。这一对应关系使条件反射现象可转化为对线性变换器上下文行为的可检验命题,区分代数结果与实证测量。代数上,推导出卡明阻断的精确闭式解,在五种学习率下模拟误差小于10^-7。实证上,预测误差修正注意力表现出线索竞争,而基于相邻性的注意力则无此现象。单状态存在两种容量模式,检索与识别的标度指数分别为1.22和1.89,符合线性和近二次预测。全头网格中,检索误差主要由总状态大小决定,而非头间分配,表明头提供容量而非冗余。我们还证明在线索正交保留测试中,分析的单状态递归无自发恢复;使用从未出现线索的对照与训练位置内的探测,训练模型亦未发现恢复。最后引入PH-注意力,一种受佩斯-霍尔启发、具显式特征索引可塑性状态的规则,实现线索依赖的学习率,是现有令牌计算门所缺乏的。

原文摘要 · Abstract (English)

Attention is widely understood as an associative memory, but that description alone does not predict how the memory will behave. Predictive theories do exist, but in the literature on animal learning. We show that the state updates of the major linear-attention families are term-for-term identical with named models from a century of animal learning theory: linear attention implements Hebbian contiguity, DeltaNet implements Rescorla--Wagner error correction, and decay variants such as RetNet implement contiguity with a stimulus trace. This dictionary turns conditioning phenomena into testable statements about the in-context behavior of linear transformers, while distinguishing algebraic consequences from empirical measurements. Algebraically, it yields an exact closed form for Kamin blocking, verified in simulation to $<10^{-7}$ across five learning rates. Empirically, it predicts a dissociation that survives training on generic in-context association: error-correcting attention exhibits cue competition, whereas contiguity-based attention does not. A single state also has two capacity regimes, with measured scaling exponents of 1.22 for faithful retrieval and 1.89 for identification, consistent with linear and near-quadratic predictions. Across the full head grid, retrieval error is governed primarily by total state size rather than its partition across heads, indicating that heads provide capacity rather than redundant copies. We also prove no spontaneous recovery for the analyzed single-state recurrences under cue-orthogonal retention trials; with a never-presented-cue control and probes within the trained positional range, we likewise find no recovery in trained models. Finally, we introduce PH-attention, a Pearce--Hall-inspired rule with an explicit feature-indexed associability state that yields cue-dependent learning rates and is absent from the token-computed gates we compare.

注意力机制学习理论线性模型认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。