arXiv:2605.12171cs.LG2026-05
证明了一层注意力无法用有理函数实现奇偶性,除非参数量线性增长。
Lower bounds for one-layer transformers that compute parity
- 用有理函数后处理的单层注意力无法计算奇偶性
- 参数量需随输入长度线性增长才能实现奇偶性
- 为注意力机制提供了理论下限,适合理论研究者
本文证明:任何经过有理函数后处理的自注意力层,若要符号表示奇偶函数,其头数与后处理函数次数的乘积必须随输入长度线性增长。结合有理函数对ReLU网络的逼近性质,进一步推导出经ReLU网络后处理的自注意力层的依赖于间隔的扩展下界。
原文摘要 · Abstract (English)
This note shows that no self-attention layer post-processed by a rational function can sign-represent the parity function unless the product of the number of heads and the degree of the post-processing function grows linearly with the input length. Combining this lower bound with rational approximation of ReLU networks yields a margin-dependent extension for self-attention layers post-processed by ReLU networks.
注意力机制理论分析奇偶性
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。