挑战视觉任务中位置编码必须保持等变性的普遍认知
A Circular Argument : Does RoPE need to be Equivariant for Vision?
- 提出非交换生成器的球面旋转位置编码(Spherical RoPE)
- 实验显示其性能与等变版本相当甚至更优
- 适合关注视觉位置编码设计效率与泛化能力的研究者
旋转位置编码(RoPE)在自然语言处理中表现优异,常被认为因其相对位置等变性而成功。本文从数学上证明,RoPE是一维数据中最通用的等变位置编码解;对于多维数据,若要求生成器可交换,则混合式RoPE为相应通用解。然而,我们质疑严格等变性对性能的关键作用。为此提出球面RoPE,基于非交换生成器,实验证明其学习行为与等变版本相当或更优。结果表明,相对位置编码并非如普遍认为般关键,尤其在计算机视觉领域。该发现有助于未来视觉位置编码研究摆脱等变性束缚,实现更快收敛与更好泛化。
原文摘要 · Abstract (English)
Rotary Positional Encodings (RoPE) have emerged as a highly effective technique for one-dimensional sequences in Natural Language Processing spurring recent progress towards generalizing RoPE to higher-dimensional data such as images and videos. The success of RoPE has been thought to be due to its positional equivariance, i.e. its status as a relative positional encoding. In this paper, we mathematically show RoPE to be one of the most general solutions for equivariant positional embedding in one-dimensional data. Moreover, we show Mixed RoPE to be the analogously general solution for M-dimensional data, if we require commutative generators -- a property necessary for RoPE's equivariance. However, we question whether strict equivariance plays a large role in RoPE's performance. We propose Spherical RoPE, a method analogous to Mixed RoPE, but assumes non-commutative generators. Empirically, we find Spherical RoPE to have the equivalent or better learning behavior compared to its equivariant analogues. This suggests that relative positional embeddings are not as important as is commonly believed, at least within computer vision. We expect this discovery to facilitate future work in positional encodings for vision that can be faster and generalize better by removing the preconception that they must be relative.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。