arXiv:2409.11995cs.LG2024-09被引 4

研究数据量增加时损失曲面如何变化,揭示收敛规律。

Unraveling the Hessian: A Key to Smooth Convergence in Loss Function Landscapes

  • 理论推导全连接网络损失曲面收敛的上界
  • 实证验证图像分类任务中损失表面随数据量收敛
  • 为确定训练样本量提供新思路,适合优化研究者

神经网络的损失景观是其训练中的关键因素,理解其特性对提升性能至关重要。本文首次系统研究了样本量增大时损失曲面的变化,理论上分析了全连接神经网络损失景观的收敛性,并推导出新增样本时损失函数值差异的上界。在多个数据集上的实证研究验证了该理论结果,表明图像分类任务中损失函数曲面随样本量增长趋于稳定。研究揭示了神经网络损失景观的局部几何特性,对样本量确定技术的发展具有重要启示。

原文摘要 · Abstract (English)

The loss landscape of neural networks is a critical aspect of their training, and understanding its properties is essential for improving their performance. In this paper, we investigate how the loss surface changes when the sample size increases, a previously unexplored issue. We theoretically analyze the convergence of the loss landscape in a fully connected neural network and derive upper bounds for the difference in loss function values when adding a new object to the sample. Our empirical study confirms these results on various datasets, demonstrating the convergence of the loss function surface for image classification tasks. Our findings provide insights into the local geometry of neural loss landscapes and have implications for the development of sample size determination techniques.

损失景观收敛性样本量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。