RoPE Transformer 自发生成类小波特征,实现多分辨率建模。
Beyond Position: the emergence of wavelet-like properties in Transformers
- 通过旋转位置编码,注意力头自发形成多分辨率处理机制。
- 该特性在训练中分阶段演化,且符合不确定性原理。
- 适用于理解大模型如何突破位置编码限制的从业者。
本文研究了采用旋转位置编码(RoPE)的Transformer模型如何自发产生类小波特性,以弥补位置编码的理论缺陷。通过对不同模型规模、架构及训练阶段的分析,我们发现注意力头会演化出类似小波变换的多分辨率处理能力。这种尺度不变的行为仅在RoPE下出现,训练过程中经历多个演化阶段,并在统计上满足基本不确定性原理。结果表明,现代Transformer的有效性源于其自发构建最优多分辨率分解的能力,用以应对固有架构约束。
原文摘要 · Abstract (English)
This paper studies how Transformer models with Rotary Position Embeddings (RoPE) develop emergent, wavelet-like properties that compensate for the positional encoding's theoretical limitations. Through an analysis spanning model scales, architectures, and training checkpoints, we show that attention heads evolve to implement multi-resolution processing analogous to wavelet transforms. We demonstrate that this scale-invariant behavior is unique to RoPE, emerges through distinct evolutionary phases during training, and statistically adheres to the fundamental uncertainty principle. Our findings suggest that the effectiveness of modern Transformers stems from their remarkable ability to spontaneously develop optimal, multi-resolution decompositions to address inherent architectural constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。