通过对称性简化,让Transformer更高效地处理关系结构。
Toward Manifest Relationality in Transformers via Symmetry Reduction
- 用不变的关联量重写表示与注意力机制
- 从源头消除模型空间和头空间的冗余自由度
- 适合关注模型效率与几何解释的研究者
Transformer模型内部存在显著冗余,源于坐标依赖的表示以及模型空间和头空间中的连续对称性。现有方法通过显式打破对称性来缓解此问题,而本文提出一种互补框架——对称性简化。我们从不变的关联量角度重新构建表示、注意力机制和优化动态,通过构造方式消除冗余自由度。该视角催生了直接作用于关系结构的架构,为减少参数冗余和分析优化过程提供了严谨的几何框架。
原文摘要 · Abstract (English)
Transformer models contain substantial internal redundancy arising from coordinate-dependent representations and continuous symmetries, in model space and in head space, respectively. While recent approaches address this by explicitly breaking symmetry, we propose a complementary framework based on symmetry reduction. We reformulate representations, attention mechanisms, and optimization dynamics in terms of invariant relational quantities, eliminating redundant degrees of freedom by construction. This perspective yields architectures that operate directly on relational structures, providing a principled geometric framework for reducing parameter redundancy and analyzing optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。