将Transformer注意力机制与多项式回归联系起来,揭示其表征演化本质。
Deriving Transformer Architectures as Implicit Multinomial Regression
- 从多项式回归视角推导注意力机制的数学基础
- 优化隐变量特征可得到与注意力一致的表征演化轨迹
- 为Transformer设计提供理论解释,适合研究者参考
尽管注意力机制在实践中被证明能提升模型性能,但缺乏严格的数学依据。本文建立了一种新颖的连接:注意力机制与多项式回归之间的关系。具体而言,在固定的多项式回归设定下,对隐变量特征进行优化所得到的解,与注意力模块对特征施加的动态演化过程一致。换言之,Transformer中表征的演化路径,可被理解为一种恢复分类最优特征的轨迹。
原文摘要 · Abstract (English)
While attention has been empirically shown to improve model performance, it lacks a rigorous mathematical justification. This short paper establishes a novel connection between attention mechanisms and multinomial regression. Specifically, we show that in a fixed multinomial regression setting, optimizing over latent features yields solutions that align with the dynamics induced on features by attention blocks. In other words, the evolution of representations through a transformer can be interpreted as a trajectory that recovers the optimal features for classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。