arXiv:2608.07323cs.LG2026-08

用新结构证明大模型前馈层无需开放正尾也能保持性能

Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU

论文配图:Is SwiGLU's Open Positive Tail Necessary? Evidence from Closed-Tail Gating with MemGLU
图 1 · 摘自论文原文
  • 提出闭尾结构MemGLU,替代SwiGLU的开放正尾设计
  • 在9M与30M参数规模下,性能仅比SwiGLU低0.1%以内
  • 揭示模型会适应预训练时的门控结构,适合架构优化研究者

我们检验解码器型语言模型的前馈网络是否需要SwiGLU的开放正尾。引入基于忆阻器分支结构的闭尾对比模型MemGLU。在三组种子、9M与30M参数量的预训练实验中,MemGLU的验证负对数似然(NLL)与SwiGLU相差不超过0.1%。训练好的SwiGLU检查点对正尾抑制敏感,机制诊断显示两者虽损失相似,但门控使用方式不同。结果表明,模型会适应预训练期间可用的门控结构。在当前测试规模下,解码器型语言模型的前馈网络无需开放正尾。

原文摘要 · Abstract (English)

We test whether decoder-only language-model FFNs require SwiGLU's open positive tail. We introduce MemGLU as a closed-tail comparator derived from a memristive branch geometry. Across paired 9M and 30M pretraining runs with three seeds, MemGLU remains within about 0.1% of SwiGLU in validation NLL. Trained SwiGLU checkpoints are sensitive to positive-tail suppression, while mechanism diagnostics show that the two models use their gates differently despite similar losses. These results suggest that models adapt to the gate geometry available during pretraining. At the tested scales, SwiGLU's open positive tail is not necessary for decoder-only language-model FFNs.

前馈网络模型架构神经网络优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。