梳理2025-2026年优化器新范式,揭示矩阵与系统级设计的演进。
Blog: Survey of Optimizers
- 按时间估计、更新几何、训练周期管理、表示与系统四维度分类优化器
- 矩阵感知方法优于传统方法,但无通用替代AdamW的方案
- 适合研究者理解优化器设计逻辑与评估标准
2025-2026年的神经网络优化已不再局限于新Adam变体的迭代。设计空间从坐标扩展到矩阵与层,从固定训练周期转向时序策略,从数学更新规则演变为需适应分片与低精度计算的状态表示。本综述基于四个独立维度组织近期优化器与训练优化方法:时间估计、更新几何、训练周期管理、表示与系统。涵盖Muon的谱归一化、Shampoo与SOAP的历史矩阵统计、自适应与混合矩阵方法、内存高效优化器、无调度训练、小批量修正及量化优化器状态。核心实证结论强调:矩阵感知方法具有实质性进步,但不存在超越上下文的AdamW替代品。性能排名随模型规模、数据-参数比、批大小、调度策略、参数划分、调参预算以及目标指标(如词元数、浮点运算量、运行时间或内存)而变化。实际影响是推动优化器设计走向组合化,并建立更严格的评估规范。
原文摘要 · Abstract (English)
Neural-network optimization in 2025-2026 is no longer well described as a succession of new Adam variants. The design space has expanded from coordinates to matrices and layers, from fixed training horizons to policies over time, and from mathematical update rules to state representations that must survive sharding and low-precision computation. This survey organizes recent optimizers and training optimization methods along four largely independent axes: temporal estimation, update geometry, horizon management, and representation and systems. It connects the spectral normalization of Muon, the historical matrix statistics of Shampoo and SOAP, adaptive and hybrid matrix methods, memory-efficient optimizers, schedule-free training, small-batch corrections, and quantized optimizer states. The central empirical conclusion is deliberately non-triumphal: matrix-aware methods represent a genuine advance, but there is no context-independent replacement for AdamW. Rankings change with model scale, data-to-parameter ratio, batch size, schedule, parameter partition, tuning budget, and whether the target metric is tokens, FLOPs, wall-clock time, or memory. The practical consequence is a compositional view of optimizer design and a stricter protocol for evaluating optimizer claims.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。