arXiv:2602.01105cs.LGcs.AI2026-02被引 1

新优化器融合谱与坐标控制,提升大模型训练效果。

OLion: Approaching the Hadamard Ideal by Intersecting Spectral and $\ell_{\infty}$ Implicit Biases

  • 结合正交更新方向与符号坐标控制,逼近海达马克约束交集。
  • 在语言与视觉任务中,性能媲美或超越AdamW和Muon。
  • 仅需动量级状态,适合大模型微调与优化器不匹配场景。

许多优化器可视为范数诱导几何下的最速下降法,因而具有相应的隐式偏差。本文提出 ameA{}( ullname{}),将正交化更新方向的谱控制与符号更新的 $\ ext{l}_ ext{\infty}$ 式坐标控制相结合。 ameA{} 构建类似 Lion 的动量方向,通过少量 Newton--Schulz 迭代近似正交化,并应用逐元素符号操作,高效逼近在谱与 $\ ext{l}_ ext{\infty}$ 约束集交集上的最大步长(矩阵参数的缩放海达马克类集合)。尽管正交化与符号操作具有强非线性,但在一个弱且经实证验证的对角各向同性假设下,我们证明了收敛性。在大规模语言与视觉训练中,包括 GPT-2、Llama 预训练、SiT 图像预训练及监督微调, ameA{} 在相近调参条件下表现优于或媲美 AdamW 与 Muon,且仅需动量级优化器状态,有效缓解了使用 AdamW 预训练检查点进行微调时的优化器不匹配问题。

原文摘要 · Abstract (English)

Many optimizers can be interpreted as steepest-descent methods under norm-induced geometries, and thus inherit corresponding implicit biases. We introduce \nameA{} (\fullname{}), which combines spectral control from orthogonalized update directions with $\ell_\infty$-style coordinate control from sign updates. \nameA{} forms a Lion-style momentum direction, approximately orthogonalizes it via a few Newton--Schulz iterations, and then applies an entrywise sign, providing an efficient approximation to taking a maximal step over the intersection of the spectral and $\ell_\infty$ constraint sets (a scaled Hadamard-like set for matrix parameters). Despite the strong nonlinearity of orthogonalization and sign, we prove convergence under a mild, empirically verified diagonal-isotropy assumption. Across large-scale language and vision training, including GPT-2 and Llama pretraining, SiT image pretraining, and supervised fine-tuning, \nameA{} matches or outperforms AdamW and Muon under comparable tuning while using only momentum-level optimizer state, and it mitigates optimizer mismatch when fine-tuning AdamW-pretrained checkpoints.

优化器大模型训练深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。