用物理启发方法训练小模型解决长序列推理,突破传统瓶颈。
Energy-Entropy Regularization: The True Power of Minimal Looped Transformers
- 引入泰利斯熵与哈密顿动力学重塑损失曲面
- 成功训练8维单头环形Transformer处理1000标记序列
- 揭示了环形注意力提升推理能力的内在机制
近期研究表明,环形Transformer相比标准深层架构具有更优的推理能力。然而,在基准任务上训练单头环形架构时,常因损失函数高度非凸且不规则而失败或表现不佳,优化过程易陷入劣质局部极小值和鞍点,难以找到全局最优解。此类单头环形Transformer的内部机制仍不清晰,从零开始训练极具挑战。本文提出一种新颖的训练框架,利用泰利斯熵与哈密顿动力学改变损失曲面的几何结构。通过将参数更新视为物理流,我们成功训练出模型维度 $d = 8$ 的单头环形Transformer,在输入序列长度为1000标记的归纳头任务中取得成功。这一成果揭示了环形结构优越推理能力的内在机理。
原文摘要 · Abstract (English)
Recent research suggests that looped Transformers have superior reasoning capabilities compared to standard deep architectures. Current approaches to training single-head looped architectures on benchmark tasks frequently fail or yield suboptimal performance due to a highly non-convex and irregular loss landscape. In these settings, optimization often stagnates in poor local minima and saddle points of the loss landscape, preventing the model from discovering the global minimum point. The internal mechanisms of these single-head looped transformer models remain poorly understood, and training them from scratch remains a significant challenge. In this paper, we propose a novel training framework that leverages Tsallis entropy and Hamiltonian dynamics to transform the geometry of the loss landscape. By treating the parameter updates as a physical flow, we successfully trained a single-head looped Transformer with model dimension $d = 8$ to solve induction head task with input sequence length of 1000 tokens. This success reveals the internal mechanism behind the superior reasoning capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。