揭示了基于旋转位置编码的张量注意力模型的理论表达局限。
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
- 从电路复杂性角度分析张量注意力与旋转位置编码的理论极限。
- 在多项式精度、常数深度等条件下,无法解决特定成员判定问题。
- 为构建更坚实的注意力机制提供理论依据,适合研究者参考。
张量注意力通过捕捉多模态间的高阶相关性,突破了传统矩阵注意力的局限。同时,旋转位置编码(RoPE)在长序列场景中表现出色,显著提升了Transformer的表达能力。然而,这些技术的理论边界仍不明确。本文分析了张量注意力及基于RoPE的张量注意力的电路复杂性,证明在多项式精度、常数深度层、线性或次线性隐藏维度下,若假设TC⁰ ≠ NC¹,它们无法解决固定成员判定问题或(A_{F,r})^*闭包问题。该发现揭示了其经验性能与理论限制之间的差距,为更理论严谨的Transformer设计与扩展提供了洞见。
原文摘要 · Abstract (English)
Tensor Attention extends traditional attention mechanisms by capturing high-order correlations across multiple modalities, addressing the limitations of classical matrix-based attention. Meanwhile, Rotary Position Embedding ($\mathsf{RoPE}$) has shown superior performance in encoding positional information in long-context scenarios, significantly enhancing transformer models' expressiveness. Despite these empirical successes, the theoretical limitations of these technologies remain underexplored. In this study, we analyze the circuit complexity of Tensor Attention and $\mathsf{RoPE}$-based Tensor Attention, showing that with polynomial precision, constant-depth layers, and linear or sublinear hidden dimension, they cannot solve fixed membership problems or $(A_{F,r})^*$ closure problems, under the assumption that $\mathsf{TC}^0 \neq \mathsf{NC}^1$. These findings highlight a gap between the empirical performance and theoretical constraints of Tensor Attention and $\mathsf{RoPE}$-based Tensor Attention Transformers, offering insights that could guide the development of more theoretically grounded approaches to Transformer model design and scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。