首次从数学上证明了Transformer训练中语法到语义的分阶段学习机制。
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
- 基于解耦的特征结构,构建简化模型分析注意力机制动态
- 发现训练过程存在语法正确后语义优化的两阶段现象
- 适用于理解语言、蛋白质等具有层次结构的数据
Transformer在真实训练中可能表现出两阶段动态。例如,在Counterfact数据集上训练GPT-2时,答案依次从句法错误、句法正确到语义正确。现有理论难以解释这种特征层面的两阶段现象,其背后或源于语法与语义等解耦特征类型。本文在包含归一化ReLU自注意力和结构化数据的简化设定下,理论上揭示了此类解耦特征结构如何引发两阶段学习动态。该解耦特征结构在实践中普遍成立,如自然语言含语法与语义,蛋白质含一级与二级结构。据我们所知,这是首个在该理论框架下对Transformer特征级两阶段优化过程的严格证明。进一步推论表明,此两阶段过程与注意力权重的谱特性密切相关。
原文摘要 · Abstract (English)
Transformers may exhibit two-stage training dynamics during the real-world training process. For instance, when training GPT-2 on the Counterfact dataset, the answers progress from syntactically incorrect to syntactically correct to semantically correct. However, existing theoretical analyses hardly account for this feature-level two-stage phenomenon, which could be conceptually attributed to disentangled two-type features like syntax and semantics. In this paper, we theoretically demonstrate how the two-stage training dynamics potentially occur in transformers. Specifically, we analyze the feature learning dynamics induced by the aforementioned disentangled two-type feature structure, grounding our analysis in a simplified yet illustrative setting that comprises normalized ReLU self-attention and structured data. Such disentanglement of feature structure is general in practice, e.g., natural languages contain syntax and semantics, and proteins contain primary and secondary structures. To our best knowledge, this is the first rigorous result regarding a feature-level two-stage optimization process in transformers within this theoretical framework. A corollary further indicates that such a two-stage process is closely related to the spectral properties of attention weights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。