arXiv:2409.14174cs.LGmath.ST2024-09被引 1

用深度网络组件构建新训练法,提升泛化能力并降低训练成本

Component-based Sketching for Deep ReLU Nets

  • 基于深度网络组件构造拟合基,将训练转为线性优化问题
  • 理论证明逼近饱和函数接近最优,泛化误差界也近乎最优
  • 相比传统梯度方法,泛化性能更优且训练开销更低

深度学习在数据挖掘与人工智能领域影响深远,但其优化与泛化之间存在矛盾:泛化良好需参数量较少的网络,而梯度算法有效收敛则依赖参数量较多的网络。为此,本文提出一种基于深度网络组件的新型结构化采样(sketching)方案。具体而言,利用具有特定效用的深度网络组件构建拟合基,体现深度网络的优势;进而将深度网络训练转化为基于该基的线性经验风险最小化问题,避免了复杂迭代算法的收敛性分析。理论分析表明,该方法对浅层网络逼近饱和函数可达到几乎最优的逼近率,并实现几乎最优的泛化误差界。数值实验显示,相较于现有梯度训练方法,该方法在保持较低训练成本的同时,展现出更优的泛化性能。

原文摘要 · Abstract (English)

Deep learning has made profound impacts in the domains of data mining and AI, distinguished by the groundbreaking achievements in numerous real-world applications and the innovative algorithm design philosophy. However, it suffers from the inconsistency issue between optimization and generalization, as achieving good generalization, guided by the bias-variance trade-off principle, favors under-parameterized networks, whereas ensuring effective convergence of gradient-based algorithms demands over-parameterized networks. To address this issue, we develop a novel sketching scheme based on deep net components for various tasks. Specifically, we use deep net components with specific efficacy to build a sketching basis that embodies the advantages of deep networks. Subsequently, we transform deep net training into a linear empirical risk minimization problem based on the constructed basis, successfully avoiding the complicated convergence analysis of iterative algorithms. The efficacy of the proposed component-based sketching is validated through both theoretical analysis and numerical experiments. Theoretically, we show that the proposed component-based sketching provides almost optimal rates in approximating saturated functions for shallow nets and also achieves almost optimal generalization error bounds. Numerically, we demonstrate that, compared with the existing gradient-based training methods, component-based sketching possesses superior generalization performance with reduced training costs.

深度网络泛化优化线性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。