提出三层神经网络在奇异参数点上的学习系数上界公式,可推广至多种激活函数。
Upper Bounds for Local Learning Coefficients of Three-Layer Neural Networks
- 基于预算、需求与供给约束的计数规则推导上界
- 在单输入情况下结果与已知精确值一致,验证了有效性
- 适用于swish等非多项式激活函数,扩展了以往适用范围
三层神经网络属于奇异性学习模型,其贝叶斯渐近行为由学习系数(即实对数规范阈值)决定。尽管该量在常规模型及部分特殊奇异模型中已被明确,但针对神经网络的通用评估方法仍有限。近期提出的半正则模型局部学习系数公式虽给出上界,但仅适用于实现参数中的非奇异点,无法用于奇异点。尤其对三层网络而言,此前上界与已知学习系数存在显著差异。本文推导出一类奇异实现参数下三层网络局部学习系数的上界公式,该公式可解释为带约束的计数规则。在非多项式实解析情形下普遍适用;在多项式情形下需满足真实分布无隐层单元的条件。本结果涵盖swish等激活函数,并在输入维数为1时,右端数值与已有精确结果一致,提供有效对比。此外,该结果系统揭示了三层数学网络权重参数对学习系数的影响机制。
原文摘要 · Abstract (English)
Three-layer neural networks are known to form singular learning models, and their Bayesian asymptotic behavior is governed by the learning coefficient, or real log canonical threshold. Although this quantity has been clarified for regular models and for some special singular models, broadly applicable methods for evaluating it in neural networks remain limited. Recently, a formula for the local learning coefficient of semiregular models was proposed, yielding an upper bound on the learning coefficient. However, this formula applies only to nonsingular points in the set of realization parameters and cannot be used at singular points. In particular, for three-layer neural networks, the resulting upper bound has been shown to differ substantially from learning coefficient values already known in some cases. In this paper, we derive a formula for an upper bound on local learning coefficients at a class of singular realization parameters in three-layer neural networks. This formula can be interpreted as a counting rule under budget, demand, and supply constraints. In the non-polynomial real-analytic case, the formula applies in general settings, whereas in the polynomial case it applies under the restriction that the true distribution has no hidden units. In particular, our result covers activation functions such as the swish function and also includes polynomial activation functions under the above restriction, thereby extending previous results to a broader class of activation functions. We further show that, when the input dimension is one, the numerical value given by the right-hand side of our upper-bound formula agrees with the previously known learning coefficient, thereby providing a useful comparison with known exact results. Our result also provides a systematic perspective on how the weight parameters of three-layer neural networks affect the learning coefficient.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。