arXiv:2607.11875cs.LGcs.AI2026-07

揭示Transformer在归纳推理中的学习机制,发现其动态可压缩到低维不变流形。

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

论文配图:Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks
图 1 · 摘自论文原文
  • 提出统一的归纳任务框架,证明注意力模型训练动态受限于低维不变流形。
  • 发现学习过程由少数可解释坐标描述,而非百万级参数,分析更简洁。
  • 可自动检测模型中已学习的电路,适用于理解复杂推理机制的科研人员。

我们提出了一个理论框架,解释了变换器语言模型中归纳推理能力的出现。以往关于变换器学习动态的研究多局限于特定任务,而本文研究了一类广义的归纳任务,该类任务统一了文献中多个合成任务,包括上下文n-gram和多跳推理。在这一类任务中,我们理论上证明了注意力模型的训练动态可被限制在一个高度可解释的、低维的不变流形上。在此流形上,学习动态由少数可解释的坐标刻画,而非数百万个参数,使理论与实证分析更为可行。利用该框架,我们揭示了数据统计如何决定上下文学习与权重学习之间的竞争,研究了随机初始化如何在多种解存在时决定‘获胜’电路,并展示了与流形相关的坐标系可用于自动检测训练后模型中已学习的电路。将电路形成视为低维动态现象,我们朝着预测性理论迈进一步。

原文摘要 · Abstract (English)

We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning. In this class, we theoretically prove that the training dynamics of attention models can be confined to a highly interpretable, low-dimensional invariant manifold. On this manifold, the learning dynamics are captured by a handful of interpretable coordinates rather than millions of parameters, making both theoretical and empirical analysis more tractable. Using this framework, we characterize how data statistics govern the competition between in-context and in-weights learning, we study how random initializations determine the `winning' circuit when multiple solutions are possible, and we demonstrate that the coordinate frame associated with the manifold can be used to automatically detect which circuits have been learned in trained models. By casting circuit formation as a low-dimensional dynamical phenomenon, we take a step toward a predictive theory of how Transformers learn.

Transformer归纳推理学习动态不变流形

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。