arXiv:2608.08922cond-mat.dis-nncond-mat.stat-mech2026-08

揭示自注意力中集群吸引子与动态凝聚的形成机制

Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention

论文配图:Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention
图 1 · 摘自论文原文
  • 通过最小化归一化自注意力动力学,发现重叠间隙决定吸引子结构
  • 当簇内相似性显著高于簇间时,簇间注意力随维度指数衰减
  • 在注意力锐度阈值以上,无序状态会自发凝聚成集群结构

Transformer 层生成依赖状态的交互网络:标记表示决定注意力矩阵,而注意力矩阵又反过来更新表示。我们研究了一个最小化的归一化自注意力动力学,在热力学极限下识别出重叠间隙是控制其吸引子结构的核心量。当标记形成内部对齐的簇,且同一簇内的相似性显著高于与其他簇的相似性时,随着维度增加,簇间注意力呈指数级抑制。该机制产生高维的集群固定点流形,从少数宏观簇到广泛微观分裂不等,并控制其对扰动的稳定性。从无序的高斯初始状态出发,我们发现只有当注意力锐度超过有限阈值时,集群态才会从弥散背景中成核,引发动态注意力凝聚相变。

原文摘要 · Abstract (English)

Transformer layers generate state-dependent interaction networks: token representations determine the attention matrix, which in turn updates the representations. We study this feedback in a minimal normalized self-attention dynamics and identify the overlap gap as the central quantity governing its attractor structure in the thermodynamic limit. When tokens form internally aligned clusters and their similarity to members of the same cluster exceeds that to every other cluster by a nonvanishing amount, inter-cluster attention is exponentially suppressed as the dimension increases. This mechanism produces a high-dimensional manifold of clustered fixed points, ranging from a few macroscopic clusters to extensive microscopic fragmentation, and also controls their stability against perturbations. Starting from an unstructured Gaussian state, we find that clustered states nucleate from the diffuse background only above a finite threshold in attention sharpness, giving rise to a dynamical attention-condensation transition.

Transformer自注意力动力系统凝聚相变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。