arXiv:2411.04990cs.LGcs.AI2024-11NeurIPS被引 37

改进注意力机制,证明任意键值矩阵下模型可收敛到单一聚类。

Clustering in Causal Attention Masking

  • 提出因果掩码下的新型注意力动态模型
  • 证明任意键-查询矩阵下渐近收敛至单个聚类
  • 首次关联经典瑞尼停车问题,探索准稳定态存在性

本文对 Geshkovski 等人(arXiv:2312.10794)提出的自注意力动力学进行了修改,以更准确反映生成式 AI 中实际使用的因果掩码注意力机制。该修改转化为一个无法解释为均场梯度流的相互作用粒子系统。尽管失去了原有结构,但我们在该背景下显著强化了 Geshkovski 等人的结果:此前严格结果仅适用于键、查询、值三矩阵均为缩放单位阵的情形,而本文证明在任意键-查询矩阵和值矩阵为单位阵时,仍可实现渐近收敛至单个聚类。此外,我们建立与组合几何中经典瑞尼停车问题的联系,为证明准稳定态的存在性迈出初步理论步伐。

原文摘要 · Abstract (English)

This work presents a modification of the self-attention dynamics proposed by Geshkovski et al. (arXiv:2312.10794) to better reflect the practically relevant, causally masked attention used in transformer architectures for generative AI. This modification translates into an interacting particle system that cannot be interpreted as a mean-field gradient flow. Despite this loss of structure, we significantly strengthen the results of Geshkovski et al. (arXiv:2312.10794) in this context: While previous rigorous results focused on cases where all three matrices (Key, Query, and Value) were scaled identities, we prove asymptotic convergence to a single cluster for arbitrary key-query matrices and a value matrix equal to the identity. Additionally, we establish a connection to the classical Rényi parking problem from combinatorial geometry to make initial theoretical steps towards demonstrating the existence of meta-stable states.

注意力机制聚类分析理论深度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。