arXiv:2410.02132cs.LGcs.NA2024-10被引 3

用函数导数信息设计非均匀初始化,提升神经网络拟合效果。

Nonuniform random feature models using derivative information

  • 根据目标函数导数设计参数分布,聚焦有效神经元区域。
  • 在Heaviside、ReLU及其光滑近似下性能优于传统均匀初始化。
  • 基于近似导数数据简化分布,采样高效且接近最优网络表现。

我们提出一种基于待逼近函数导数信息的非均匀数据驱动参数分布,用于浅层神经网络初始化。该方法建立在非参数回归框架下,相比传统的均匀随机特征模型具有明显优势。研究涵盖Heaviside与ReLU激活函数及其光滑近似(sigmoid与softplus),并利用全训练最优网络产生的谐波分析与稀疏表示最新成果。通过扩展精确表示的解析结果,得到集中在对应局部导数建模能力较强的参数空间区域的密度函数。基于此,我们提出基于输入点近似导数数据的密度简化方案,实现高效采样,在多种场景下使随机特征模型性能接近最优网络。

原文摘要 · Abstract (English)

We propose nonuniform data-driven parameter distributions for neural network initialization based on derivative data of the function to be approximated. These parameter distributions are developed in the context of non-parametric regression models based on shallow neural networks, and compare favorably to well-established uniform random feature models based on conventional weight initialization. We address the cases of Heaviside and ReLU activation functions, and their smooth approximations (sigmoid and softplus), and use recent results on the harmonic analysis and sparse representation of neural networks resulting from fully trained optimal networks. Extending analytic results that give exact representation, we obtain densities that concentrate in regions of the parameter space corresponding to neurons that are well suited to model the local derivatives of the unknown function. Based on these results, we suggest simplifications of these exact densities based on approximate derivative data in the input points that allow for very efficient sampling and lead to performance of random feature models close to optimal networks in several scenarios.

神经网络初始化随机特征导数信息非均匀分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。