分析深度线性与ReLU网络的学习效率,揭示奇异学习模型的泛化能力机制。
Singular leaning coefficients and efficiency in learning theory
- 通过学习系数衡量深度模型泛化效率,构建理论分析框架。
- 首次计算三层ReLU网络与深度线性模型的学习系数,揭示其收敛规律。
- 拓展至Softmax模型,为复杂神经网络提供理论支持,适合理论研究者。
具有非正定费舍尔信息矩阵的奇异学习模型,包括神经网络、降秩回归、玻尔兹曼机、正态混合模型等,在学习机器的发展中被广泛应用。然而,其理论分析仍处于初期阶段。本文研究了学习系数,该系数表征深度线性学习模型及含ReLU单元的三层神经网络模型的通用学习效率。最后,将结果推广至Softmax函数的情形。
原文摘要 · Abstract (English)
Singular learning models with non-positive Fisher information matrices include neural networks, reduced-rank regression, Boltzmann machines, normal mixture models, and others. These models have been widely used in the development of learning machines. However, theoretical analysis is still in its early stages. In this paper, we examine learning coefficients, which indicate the general learning efficiency of deep linear learning models and three-layer neural network models with ReLU units. Finally, we extend the results to include the case of the Softmax function.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。