动态损失函数通过周期性调整类别权重,改善神经网络学习效果
Dynamical loss functions shape landscape topography and improve learning in artificial neural networks
- 用周期性变化的损失权重重构标准损失函数
- 在不同规模网络上显著提升验证准确率
- 揭示训练中景观不稳定性与边缘稳定性最小化的关联
动态损失函数源自监督分类中的标准损失函数,但对每类贡献进行周期性增减调整。这种振荡全局改变损失景观,却不影响全局最小值。本文展示如何将交叉熵和均方误差转化为动态损失函数。首先分析网络规模或学习率增大对极小值深度与尖锐度的影响,基于此提出多种动态损失函数版本,并在简单分类任务中验证其可显著提升不同规模网络的验证准确率。最后研究训练过程中动态损失函数景观的演化,揭示可能与边缘稳定性最小化相关的不稳定性出现。
原文摘要 · Abstract (English)
Dynamical loss functions are derived from standard loss functions used in supervised classification tasks, but are modified so that the contribution from each class periodically increases and decreases. These oscillations globally alter the loss landscape without affecting the global minima. In this paper, we demonstrate how to transform cross-entropy and mean squared error into dynamical loss functions. We begin by discussing the impact of increasing the size of the neural network or the learning rate on the depth and sharpness of the minima that the system explores. Building on this intuition, we propose several versions of dynamical loss functions and use a simple classification problem where we can show how they significantly improve validation accuracy for networks of varying sizes. Finally, we explore how the landscape of these dynamical loss functions evolves during training, highlighting the emergence of instabilities that may be linked to edge-of-instability minimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。