对比多种优化算法在扩散模型训练中的表现,发现新方法可降低18%最终损失。
Optimization Benchmark for Diffusion Models on Dynamical Systems
- 采用最新优化算法训练扩散模型以去噪轨迹
- Muon和SOAP比AdamW最终损失低18%
- 适用于希望提升训练效率的机器学习研究者
扩散模型的训练常被忽视在新优化技术评估中。本文针对去噪轨迹的扩散模型训练,基准测试了近期优化算法。结果显示,Muon与SOAP相比AdamW可实现18%更低的最终损失。同时,重新审视了文本或图像模型训练中若干近期现象在扩散模型训练中的影响,包括学习率调度对训练动态的作用,以及Adam与SGD之间的性能差距。
原文摘要 · Abstract (English)
The training of diffusion models is often absent in the evaluation of new optimization techniques. In this work, we benchmark recent optimization algorithms for training a diffusion model for denoising flow trajectories. We observe that Muon and SOAP are highly efficient alternatives to AdamW (18% lower final loss). We also revisit several recent phenomena related to the training of models for text or image applications in the context of diffusion model training. This includes the impact of the learning-rate schedule on the training dynamics, and the performance gap between Adam and SGD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。