arXiv:2604.09560cs.LG2026-04

将注意力机制与扩散过程统一为几何运算,揭示其内在的度量、势能和环流结构。

The Diffusion-Attention Connection

  • 将注意力视为扩散映射的行归一化,分解出度量、势能与环流三类几何分量
  • 正半定几何约束几乎无代价,而移除环流在归纳任务中导致显著性能下降
  • 为注意力头提供可测量的几何解释,适用于理解模型内部机制的研究者

Softmax注意力是扩散映射的行归一化算子:两者均将学习到的得分转化为马尔可夫算子,仅在允许的得分内容上有所不同。分解该得分揭示三个几何部分:具有Witten-Laplacian连续极限的度量核心,对应马尔可夫-沃滕测度变换的精确节点势能部分,以及作为不可逆马尔可夫-Girsanov传输或磁U(1)相位实现的环流部分。注意力因此可在熟悉的数学框架下被审计:每个训练后的注意力头都携带可度量的几何、势能与通量。标准机制也获得几何定位——Coifman-Lafon归一化即精确密度修正,旋转嵌入为纯规范场,AdaLN为Cauchy-Green形变结合Witten形变。在预训练扩散变换器与语言模型上的实验验证了这一分解:强制正半定几何几乎无代价,与此识别一致;而移除环流则带来显著代价,尤其在归纳任务中表现最差。

原文摘要 · Abstract (English)

Softmax attention is the row-normalized operator of a diffusion map: both normalize a learned score into a Markov operator, and differ only in what the score is allowed to contain. Decomposing that score reveals three geometric sectors: a metric core with Witten-Laplacian continuum limit, an exact node-potential sector corresponding to a Markov--Witten change of measure, and a circulating sector realized as irreversible Markov--Girsanov transport or as a magnetic $U(1)$ phase. Attention thereby becomes auditable in familiar mathematics: every trained head carries measurable geometry, potential, and flux, while standard mechanisms acquire geometric addresses---Coifman--Lafon normalization as an exact density correction, rotary embeddings as pure gauge, and AdaLN as a Cauchy--Green deformation combined with an Witten deformation. Experiments on pretrained diffusion transformers and language models test this decomposition: enforcing positive-semi-definite geometry is nearly free, consistent with the identification, whereas removing circulation incurs a substantial cost, sharpest on induction.

注意力机制扩散模型几何分析可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。