发现Transformer训练边界具有分形结构,揭示了超参数敏感性的深层机制。
Mapping the Edge of Chaos: Fractal-Like Boundaries in The Trainability of Decoder-Only Transformer Models
- 通过分析学习率空间,发现训练稳定性边界呈现自相似分形结构。
- 在多尺度下观测到稳定与发散区域间存在重复的复杂边界模式。
- 适合研究模型训练鲁棒性或超参数优化的科研人员参考。
在分形几何中,简单的迭代过程可生成复杂的结构,将参数空间划分为稳定与不稳定区域。类似地,训练大语言模型时,即使微调超参数(如使用Adam优化器),也可能使训练从收敛变为发散。近期对小型神经网络的研究表明,这种状态转变的边界具有分形特征。本研究将该发现拓展至中等规模的解码器仅型Transformer架构,采用更一致的收敛度量,并考察注意力层与全连接层的学习率超参数空间。结果表明,训练能力的边界并非简单阈值,而是在多个尺度上形成自相似但看似随机的结构,具有统计上一致且重复的模式。在稳定收敛区域周围,存在一个复杂的混沌边界,体现了底层训练动力学的高度敏感性。
原文摘要 · Abstract (English)
In the realm of fractal geometry, intricate structures emerge from simple iterative processes that partition parameter spaces into regions of stability and instability. Likewise, training large language models involves iteratively applying update functions, such as Adam, where even slight hyperparameter adjustments can shift the training process from convergence to divergence. Recent evidence from miniature neural networks suggests that the boundary separating these outcomes displays fractal characteristics. Building on these insights, this study extends them to medium-sized, decoder-only transformer architectures by employing a more consistent convergence measure and examining the learning rate hyperparameter landscape for attention and fully connected layers. The results show that the trainability frontier is not a simple threshold; rather, it forms a self-similar yet seemingly random structure at multiple scales, with statistically consistent and repeating patterns. Within this landscape, a region of stable convergence is surrounded by a complex chaotic border, illustrating the sensitive nature of the underlying training dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。