arXiv:2608.11909cs.LGcs.FL2026-08

揭示旋转位置编码的表达能力本质,区分理论与实际表现差异

Disentangling the Expressivity of RoPE

  • 从形式化角度分析旋转位置编码的周期性与非周期性机制
  • 周期性设定可表征模运算语言,而实际编码仅模拟局部偏移
  • 适合研究位置编码理论、模型设计与长序列建模的读者

旋转位置编码(RoPE)成功的原因存在两种解释:一种是基于周期性位置信息与模运算谓词的表达能力,另一种则强调位置锚点与局部偏移。本文在全均匀、有限精度的软注意力变换器框架下形式化了这两种观点。结果表明,若所有旋转分量均为周期性,则RoPE变换器可识别恰好由带模运算的过去时态逻辑定义的语言。然而,常规RoPE的旋转永不重复,仅在精度依赖下有限地模拟固定偏移回溯算子,而非对所有长度实现模运算刻画。受控实验验证了这一差异:构造的周期性调度在模运算语言上具有长度泛化能力,而传统RoPE更像局部性偏差,可能损害需要远距离上下文不变访问的任务。整体研究为实际使用的RoPE变换器提供了理论表达能力的新理解。

原文摘要 · Abstract (English)

Two accounts recur in explanations of the success of rotary position embeddings (RoPE). Expressivity studies associate periodic position information with modular predicates, whereas mechanistic and long-context studies emphasize positional anchors and local offsets. We formalize both accounts for fully uniform, finite-precision soft-attention transformers. We find that, if every rotary component is periodic, RoPE transformers recognize exactly the languages definable in past temporal logic with modular predicates. Conventional RoPE is different: The rotations it computes never repeat. This yields a precision-dependent bounded simulation of fixed-offset look-back operators, rather than an all-length modular characterization. Controlled experiments match this separation: Constructed periodic schedules length-generalize on modular languages, while conventional RoPE behaves more like a bounded locality bias and can impair tasks requiring position-invariant access to distant context. Altogether, our findings shed light on RoPE transformers, bringing theoretical expressivity characterizations closer to models used in practice.

位置编码表达能力变换器形式化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。