arXiv:2504.18590cs.LGcs.AI2025-04
通过微分方程视角,分层调整离散化加速Transformer训练。
A multilevel approach to accelerate the training of Transformers
- 用常微分方程解析Transformer,设计分层离散策略。
- 实验表明训练速度显著提升,性能保持稳定。
- 适合追求高效训练的深度学习研究者。
本文研究了多层级方法在加速Transformer架构训练中的潜力。基于对这些架构的常微分方程(ODE)解释,我们提出了一种合理的方法,通过动态调整这些ODE Transformer的离散化程度来加速训练过程。通过与标准训练流程的实验对比,验证了该方法的有效性。
原文摘要 · Abstract (English)
In this article, we investigate the potential of multilevel approaches to accelerate the training of transformer architectures. Using an ordinary differential equation (ODE) interpretation of these architectures, we propose an appropriate way of varying the discretization of these ODE Transformers in order to accelerate the training. We validate our approach experimentally by a comparison with the standard training procedure.
Transformer训练加速微分方程
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。