模型能力突现源于稀疏注意力模式的随机学习,大模型更早掌握。
Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns

- 通过合成数据训练,发现能力突现源于任务相关注意力模式的突然习得。
- 长上下文和低稀疏度使注意力模式更难学,大模型平均提前获得能力。
- 揭示了突现机制:稀疏注意力学习难度导致能力非平滑涌现,适合研究模型本质者看。
Transformer语言模型的神经规模定律预测预训练损失随参数增加而平滑下降,但下游能力如上下文学习在特定模型规模后会突然涌现。本文表明,这些突现能力在训练过程中随机出现,且大模型平均更早获得。我们证明,模式补全和间接宾语识别等能力的涌现对应于任务相关注意力模式的突然学习。通过在合成线性映射和细胞自动机数据集上训练Transformer模型,发现注意力模式的学习难度取决于上下文长度和模式稀疏度。增加注意力头数可提升合成任务的学习效率,而增大头维度在达到最小容量后收益递减。此外,我们研究了替代注意力机制的架构,发现MLP-Mixer在具有复杂注意力模式的线性映射任务上优于Transformer。结果揭示了突现的能力机制:由于变压器模型中稀疏注意力模式的内在学习难度,下游能力会突然出现。
原文摘要 · Abstract (English)
Neural scaling laws for transformer language models predict smooth improvements in pretraining loss with increasing parameters, but downstream capabilities such as in-context learning are known to emerge abruptly past a certain model scale. In this paper, we show that emergent capabilities arise stochastically throughout training, with larger models acquiring them earlier on average. We demonstrate that the emergence of capabilities such as pattern completion and indirect object identification corresponds to the abrupt learning of task-relevant attention patterns. To isolate this phenomenon, we train transformer models on synthetic linear map and cellular automata datasets, and we show that the difficulty of learning attention patterns depends on context length and pattern sparsity. Moreover, scaling the number of attention heads improves learning efficiency on our synthetic tasks, while increasing the head dimension yields diminishing returns past a minimum capacity. We additionally investigate architectures with alternative attention mechanisms, showing that MLP-Mixer outperforms a transformer on linear map tasks with complex attention patterns. Our findings provide a mechanistic insight into emergence, showing that downstream capabilities arise abruptly due to the intrinsic difficulty of learning sparse attention patterns in transformer models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。