超叠加机制让训练速度提升十倍,且不受数据影响。
Superposition unifies power-law training dynamics
- 用师生框架分析特征叠加对训练的影响。
- 叠加导致训练指数趋近1,速度比无叠加快十倍。
- 适合研究大模型训练效率与通用性的人参考。
我们通过师生框架研究了特征叠加在幂律训练动态中的作用。首先推导出无叠加情况下的解析理论,表明幂律训练指数取决于输入数据统计和通道重要性。令人惊讶的是,叠加瓶颈会引发向统一的幂律指数∼1的转变,该指数独立于数据和通道统计。有叠加时的1/时间训练相比无叠加时的纯序列学习可实现高达十倍的加速。这一发现表明,叠加可带来与数据无关的快速训练,对采用叠加的多种神经网络(包括大规模语言模型)具有重要意义。
原文摘要 · Abstract (English)
We investigate the role of feature superposition in the emergence of power-law training dynamics using a teacher-student framework. We first derive an analytic theory for training without superposition, establishing that the power-law training exponent depends on both the input data statistics and channel importance. Remarkably, we discover that a superposition bottleneck induces a transition to a universal power-law exponent of $\sim 1$, independent of data and channel statistics. This one over time training with superposition represents an up to tenfold acceleration compared to the purely sequential learning that takes place in the absence of superposition. Our finding that superposition leads to rapid training with a data-independent power law exponent may have important implications for a wide range of neural networks that employ superposition, including production-scale large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。