arXiv:2605.09991cs.AIcs.LG2026-05

不同优化器导致的模型解空间结构差异,影响训练结果的连通性。

Optimizer-Induced Mode Connectivity: From AdamW to Muon

论文配图:Optimizer-Induced Mode Connectivity: From AdamW to Muon
图 1 · 摘自论文原文
  • 用优化器诱导的隐式正则化研究解的连通性
  • 大宽度下同优化器解可连通,小宽度时分离且有损失屏障
  • 适合关注优化器对模型结构影响的研究者

模式连通性被广泛研究,但优化器的作用仍不明确。本文通过优化器诱导的隐式正则化重新审视该问题,探讨在特定优化器约束下解的连通性。对于两层ReLU网络,我们证明在足够大的宽度下,同一优化器(如AdamW、Muon或Lion-$\mathcal{K}$族)产生的解构成连通集合,这一结果无法由以往工作推导。进一步分析不同优化器区域的交互:大宽度时区域可能分离或重叠,取决于正则化;小宽度示例中,AdamW与Muon收敛至被明确损失屏障分隔的零损失分支。在GPT-2预训练中,同优化器路径保持模型谱不变,跨优化器路径则呈现平滑过渡。结果揭示了经典模式连通性之外的优化器依赖结构。

原文摘要 · Abstract (English)

Mode connectivity has been widely studied, yet the role of the optimizer remains underexplored. We revisit it through optimizer-induced implicit regularization, asking how connectivity behaves when restricted to solutions constrained by a given optimizer. For two-layer ReLU networks, we show that solutions from a single optimizer -- AdamW, Muon, or others in the Lion-$\mathcal{K}$ family -- form a connected set at sufficiently large width, a result not implied by prior work. We then characterize how optimizer-induced regions interact: at large width two different regions can be disjoint or overlap depending on regularization, while in our small-width example AdamW and Muon converge to disconnected zero-loss components separated by a provable loss barrier. Empirically, in GPT-2 pretraining, we observe same-optimizer paths preserve each model's spectrum while cross-optimizer paths traverse a smooth transition. Our results reveal optimizer-dependent structure beyond classical mode connectivity literature.

优化器模式连通性深度学习理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。