arXiv:2512.25060cs.LG2025-12被引 3

不同注意力架构的模加法实现本质相同,几何拓扑一致。

On the geometry and topology of representations: the manifolds of modular addition

  • 通过集体分析神经元群,发现每种表示都是可研究的流形。
  • 跨数百个电路统计验证,不同架构的模加法表示高度相似。
  • 适用于研究深度学习模型内在表示结构的学者。

钟表与披萨两种解释分别对应具有固定注意力和可学习注意力的架构,曾被用来论证不同设计会产生不同的模加法电路。本文表明,这两种架构实际实现了相同的算法,其学习到的表示在拓扑与几何上等价。我们的方法超越了对单个神经元与权重的解读,转而识别每个学习表示所对应的全部神经元,并将这些神经元群体作为一个整体进行研究。这一视角揭示,每个学习表示构成一个流形,可用拓扑工具进行分析。基于此,我们对数百个自然生成的模加法电路进行了统计分析,证实了主流深度学习范式下学习到的模加法电路具有高度一致性。

原文摘要 · Abstract (English)

The Clock and Pizza interpretations, associated with architectures differing in either uniform or learnable attention, were introduced to argue that different architectural designs can yield distinct circuits for modular addition. In this work, we show that this is not the case, and that both uniform attention and trainable attention architectures implement the same algorithm via topologically and geometrically equivalent representations. Our methodology goes beyond the interpretation of individual neurons and weights. Instead, we identify all of the neurons corresponding to each learned representation and then study the collective group of neurons as one entity. This method reveals that each learned representation is a manifold that we can study utilizing tools from topology. Based on this insight, we can statistically analyze the learned representations across hundreds of circuits to demonstrate the similarity between learned modular addition circuits that arise naturally from common deep learning paradigms.

表示学习拓扑分析模加法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。