arXiv:2602.01863stat.MLcs.LG2026-02被引 3

将Transformer看作概率测度的联想记忆,给出可证明泛化性能的理论框架。

Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality

  • 用概率测度重释注意力机制,将上下文视为词元分布
  • 在谱假设下,浅层Transformer可学习召回与预测映射,收敛速率最优
  • 适用于设计可证明性能的长上下文模型,适合理论研究者

Transformers 通过内容寻址式检索和处理任意长度上下文的能力表现出色。本文从概率测度层面重构联想记忆:将上下文视为词元分布,注意力视为测度上的积分算子。对于混合上下文 $ν= I^{-1} \ sum_{i=1}^I μ^{(i^*)}$ 和查询 $x_{ {q}}(i^*)$,任务分解为(i)相关分量 $μ^{(i^*)}$ 的召回,(ii)基于 $(μ_{i^*},x_ {q})$ 的预测。研究了由经验风险最小化训练的软注意力(非固定核),并证明在输入密度满足谱假设条件下,一个浅层测度论Transformer加MLP能学习该召回-预测映射。进一步建立了具有相同指数阶的匹配极小极大下界,证明了收敛阶的紧致性。该框架为设计和分析从任意长分布上下文中回忆的Transformer提供了严谨方法,并具备可证明的泛化保证。

原文摘要 · Abstract (English)

Transformers excel through content-addressable retrieval and the ability to exploit contexts of, in principle, unbounded length. We recast associative memory at the level of probability measures, treating a context as a distribution over tokens and viewing attention as an integral operator on measures. Concretely, for mixture contexts $ν= I^{-1} \sum_{i=1}^I μ^{(i^*)}$ and a query $x_{\mathrm{q}}(i^*)$, the task decomposes into (i) recall of the relevant component $μ^{(i^*)}$ and (ii) prediction from $(μ_{i^*},x_\mathrm{q})$. We study learned softmax attention (not a frozen kernel) trained by empirical risk minimization and show that a shallow measure-theoretic Transformer composed with an MLP learns the recall-and-predict map under a spectral assumption on the input densities. We further establish a matching minimax lower bound with the same rate exponent (up to multiplicative constants), proving sharpness of the convergence order. The framework offers a principled recipe for designing and analyzing Transformers that recall from arbitrarily long, distributional contexts with provable generalization guarantees.

Transformer理论分析注意力机制泛化保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。