arXiv:2501.13296cs.LGstat.ML2025-01被引 3

提出可实时估算重要性采样方差降低的高效训练方法

Exploring Variance Reduction in Importance Sampling for Efficient DNN Training

  • 仅用重要性采样小批量估计方差减少量
  • 实测相比现有方法提升训练效率与模型准确率
  • 适合需要高效训练的深度学习研究者

重要性采样广泛用于降低深度神经网络(DNN)训练中梯度估计的方差,从而提升训练效率。然而,由于计算开销,高效评估相对于均匀采样的方差降低仍具挑战。本文提出一种仅依赖重要性采样小批量即可估计方差减少的方法,并基于此设计了自动学习率调整的有效小批量大小。同时引入一个量化重要性采样效率的绝对指标,并提出基于移动梯度统计的实时重要性评分算法。理论分析与基准数据集上的实验表明,所提算法在保持极低计算开销的前提下,持续降低方差,提升训练效率与模型准确率,优于当前主流重要性采样方法。

原文摘要 · Abstract (English)

Importance sampling is widely used to improve the efficiency of deep neural network (DNN) training by reducing the variance of gradient estimators. However, efficiently assessing the variance reduction relative to uniform sampling remains challenging due to computational overhead. This paper proposes a method for estimating variance reduction during DNN training using only minibatches sampled under importance sampling. By leveraging the proposed method, the paper also proposes an effective minibatch size to enable automatic learning rate adjustment. An absolute metric to quantify the efficiency of importance sampling is also introduced as well as an algorithm for real-time estimation of importance scores based on moving gradient statistics. Theoretical analysis and experiments on benchmark datasets demonstrated that the proposed algorithm consistently reduces variance, improves training efficiency, and enhances model accuracy compared with current importance-sampling approaches while maintaining minimal computational overhead.

深度学习重要性采样方差降低训练效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。