改进SGD优化器,融合数值方法提升神经网络训练效果。
Enhancing Deep Learning with Optimized Gradient Descent: Bridging Numerical Methods and Neural Network Training
- 借鉴数值优化方法改进SGD,增强算法可解释性。
- 在多个深度学习任务中验证了新算法的优越性能。
- 适合关注优化器设计与模型训练效率的研究者。
优化理论是实现系统最优性能的关键工具,其起源可追溯至经济领域以识别最佳投资策略。从古希腊几何到牛顿与莱布尼茨的微积分,再到拉格朗日、柯西和冯·诺依曼等人的贡献,优化理论历经数百年发展。现代计算机科学的兴起推动了其广泛应用,涵盖工程、决策分析与运筹学等领域。本文深入探讨优化理论与深度学习之间的深刻联系,强调后者中优化问题的普遍性。研究聚焦于梯度下降及其变体,这些是优化神经网络的核心方法。本文提出一种对SGD优化器的改进,灵感源自数值优化方法,旨在提升算法的可解释性与准确性。在多种深度学习任务上的实验验证了新算法的有效性。论文最后强调,优化理论持续演进,在解决复杂问题、提升计算能力及支持更优政策制定方面发挥着日益重要的作用。
原文摘要 · Abstract (English)
Optimization theory serves as a pivotal scientific instrument for achieving optimal system performance, with its origins in economic applications to identify the best investment strategies for maximizing benefits. Over the centuries, from the geometric inquiries of ancient Greece to the calculus contributions by Newton and Leibniz, optimization theory has significantly advanced. The persistent work of scientists like Lagrange, Cauchy, and von Neumann has fortified its progress. The modern era has seen an unprecedented expansion of optimization theory applications, particularly with the growth of computer science, enabling more sophisticated computational practices and widespread utilization across engineering, decision analysis, and operations research. This paper delves into the profound relationship between optimization theory and deep learning, highlighting the omnipresence of optimization problems in the latter. We explore the gradient descent algorithm and its variants, which are the cornerstone of optimizing neural networks. The chapter introduces an enhancement to the SGD optimizer, drawing inspiration from numerical optimization methods, aiming to enhance interpretability and accuracy. Our experiments on diverse deep learning tasks substantiate the improved algorithm's efficacy. The paper concludes by emphasizing the continuous development of optimization theory and its expanding role in solving intricate problems, enhancing computational capabilities, and informing better policy decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。