提出延迟投影方法,让大核模型训练更快更高效。
Fast training of large kernel models with delayed projections
- 用延迟投影改进预条件随机梯度下降,支持更大模型训练
- 在多个数据集上训练速度显著提升,精度相当或更好
- 适合需要高效大模型核学习的科研与工程人员
经典核机器在扩展到大规模数据集和模型规模时面临显著挑战——这正是神经网络成功的关键因素。本文提出一种构建可扩展核机器的新方法,该方法通过在预条件随机梯度下降(PSGD)中引入延迟投影,实现了对更大模型的高效训练,突破了核学习的实际规模限制。我们在多个数据集上验证了算法EigenPro4,结果表明其相比现有方法实现了显著的训练加速,同时保持相当或更优的分类精度。
原文摘要 · Abstract (English)
Classical kernel machines have historically faced significant challenges in scaling to large datasets and model sizes--a key ingredient that has driven the success of neural networks. In this paper, we present a new methodology for building kernel machines that can scale efficiently with both data size and model size. Our algorithm introduces delayed projections to Preconditioned Stochastic Gradient Descent (PSGD) allowing the training of much larger models than was previously feasible, pushing the practical limits of kernel-based learning. We validate our algorithm, EigenPro4, across multiple datasets, demonstrating drastic training speed up over the existing methods while maintaining comparable or better classification accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。