arXiv:2505.17863cs.LGcs.NE2025-05NeurIPS被引 26

发现稀疏注意力的涌现规律,揭示数据分布与重复训练的影响

The emergence of sparse attention: impact of data distribution and benefits of repetition

  • 通过小模型实证+理论分析,揭示稀疏注意力的形成机制
  • 发现涌现时间服从幂律,受任务结构、架构和优化器影响
  • 重复训练可显著加速注意力模式的出现,适合关注训练动态的研究者

大规模语言模型和神经网络普遍存在涌现现象:随着模型规模扩大和训练时间延长,某些能力会突然出现。尽管已有初步研究,我们对这种现象何时发生及如何产生仍缺乏全面理解。为此,本文研究了变压器(Transformers)中一种关键且常见注意力模式——稀疏注意力的训练涌现过程。结合玩具模型的理论分析与在类线性回归任务上小型Transformer的实证观察,我们揭示了驱动稀疏注意力涌现的机制,并发现其出现时机遵循基于任务结构、模型架构和优化器选择的幂律规律。此外,我们发现重复训练能极大加速该现象的发生。最后,我们在经典的上下文关联回忆任务上验证了上述结论。研究结果提供了一个简洁且理论支持充分的框架,用于理解数据分布与模型设计如何影响一种涌现行为的学习动态。

原文摘要 · Abstract (English)

Emergence is a fascinating property of large language models and neural networks more broadly: as models scale and train for longer, they sometimes develop new abilities in sudden ways. Despite initial studies, we still lack a comprehensive understanding of how and when these abilities emerge. To address this gap, we study the emergence over training of sparse attention, a critical and frequently observed attention pattern in Transformers. By combining theoretical analysis of a toy model with empirical observations on small Transformers trained on a linear regression variant, we uncover the mechanics driving sparse attention emergence and reveal that emergence timing follows power laws based on task structure, architecture, and optimizer choice. We additionally find that repetition can greatly speed up emergence. Finally, we confirm these results on a well-studied in-context associative recall task. Our findings provide a simple, theoretically grounded framework for understanding how data distributions and model design influence the learning dynamics behind one form of emergence.

注意力机制模型涌现训练动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。