arXiv:2607.07066cs.LGcs.AI2026-07被引 1

小规模Transformer学会模乘运算,揭示局部代数结构与傅里叶机制的作用。

Multiplication Beyond Groups: Stratified Fourier Mechanisms in Transformer Circuits

论文配图:Multiplication Beyond Groups: Stratified Fourier Mechanisms in Transformer Circuits
图 1 · 摘自论文原文
  • 提出单半群扩展,用局部代数区域替代全局表示空间。
  • 模型在平方自由模乘任务中表现良好,注意力具类别敏感路由与低秩写方向。
  • 适用于研究Transformer中非群运算的可解释性与代数机制。

Transformer展现出强大的算法推理能力,但现有机制分析多集中于全局可逆操作(如循环加法和群合成)。本文研究小规模Transformer如何学习复合模数下的模整数乘法,该运算因零因子存在而不可逆。我们提出单半群扩展:通过表示(GCR)对群合成的局部化推广,表明模型并非依赖单一全局表示空间,而是将输入空间划分为局部分层代数区域,其中类群结构仍存在并可应用傅里叶机制。在训练于平方自由模乘任务的Transformer中,嵌入围绕这些区域组织,注意力表现出类别敏感路由与低秩写方向,局部特征解释了模型输出逻辑的大部分差异。结果表明,先前用于群操作的表示论机制可拓展至更一般的代数结构。

原文摘要 · Abstract (English)

Transformers have demonstrated a remarkable ability to learn algorithmic reasoning, yet mechanistic analyses have mostly focused on globally invertible operations such as cyclic addition and group composition. In this work, we investigate how small transformers learn modular integer multiplication over composite moduli, a fundamentally non-invertible operation due to the presence of zero-divisors. We propose the monoid extension: a localized generalization of Group Composition via Representation (GCR) that suggests the learned computation does not rely on a single global representation space. Instead, the model partitions the input space into local hierarchical algebraic regions, where group-like structure survives and Fourier mechanisms can be applied. In transformers trained on square-free modular multiplication, we find that embeddings organize around these regions, attention exhibits class-sensitive routing and low-rank write directions, and local character features explain a large fraction of the model's output logits. Our results suggest that representation-theoretic mechanisms previously identified for group operations can extend beyond groups to more general structures.

Transformer代数结构傅里叶机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。