arXiv:2608.30417cs.LGmath.AG2026-08

固定结构的等变注意力无法覆盖所有等变映射,会导致表达力损失。

No Equivariant Architecture Covers All Equivariant Attention

  • 等变注意力需满足特定群作用下的参数约束
  • 单个架构只能覆盖极多不可约成分中的一个
  • 适合研究等变神经网络表达能力的学者

我们对等变多头自注意力(MHSA)给出了完整刻画:若某MHSA层对对称群 $G$ 等变,则 $G$ 只能通过置换头簇作用,且查询-键(QK)与输出-值(OV)矩阵需满足与群作用相关的等变约束。因此证明,任何通过多项式参数化无约束参数来实现精确等变性的固定MHSA架构,必然在等变映射类中导致表达力损失:无约束MHSA的等变轨迹在简化参数空间中构成极多的扎里斯基不可约分支,而任一架构最多覆盖其中一个。当 $G=D_4$ 作用于 $C$ 个正则表示副本作为标记特征空间时,有 $inom{C}{64}$ 个组件。

原文摘要 · Abstract (English)

We give a complete characterization of equivariant multi-head self-attention (MHSA): if an MHSA layer is equivariant to a symmetry group $G$, then $G$ can only act by permuting head-clusters, with QK and OV matrices satisfying an equivariance constraint tied to the group action. As a consequence, we prove that any fixed MHSA architecture that achieves exact equivariance by polynomially parameterizing unconstrained MHSA parameters inevitably leads to expressivity loss within the class of equivariant maps: the equivariance locus of unconstrained MHSA forms a union of extremely many Zariski-irreducible components in a reduced parameter space, and any single architecture covers at most one. For $G=D_4$ acting on $C$ copies of the regular representation as the token feature space, we show that there are $Ω(C^{64})$ components for eight attention heads.

注意力机制等变性表达力损失群作用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。