分析神经网络隐藏层特征的自相似性,发现可控自相似性可提升分类性能。
Self-similarity Analysis in Deep Neural Networks
- 构建基于隐藏层输出的复杂网络模型,量化不同层的自相似性
- 在MLP和注意力模型中,自相似性约束使性能最高提升6个百分点
- 揭示了不同架构自相似性差异,为优化训练提供新思路
现有研究发现部分深度神经网络在特征表示或参数分布上表现出强层级自相似性。然而,除初步探讨权重幂律分布对模型性能的影响外,尚缺乏对隐藏空间几何自相似性如何影响权重优化的定量分析,也未明确内部神经元的动态行为。本文提出一种基于隐藏层神经元输出特征的复杂网络建模方法,研究不同隐藏层构建的特征网络的自相似性,并分析调整特征网络自相似度对分类性能的影响。在MLP、卷积网络和注意力架构三类模型上验证,结果表明不同模型架构的特征网络自相似性程度各异。在训练过程中施加自相似性嵌入约束,可使自相似型深度神经网络(如MLP和注意力架构)的性能最高提升6个百分点。
原文摘要 · Abstract (English)
Current research has found that some deep neural networks exhibit strong hierarchical self-similarity in feature representation or parameter distribution. However, aside from preliminary studies on how the power-law distribution of weights across different training stages affects model performance,there has been no quantitative analysis on how the self-similarity of hidden space geometry influences model weight optimization, nor is there a clear understanding of the dynamic behavior of internal neurons. Therefore, this paper proposes a complex network modeling method based on the output features of hidden-layer neurons to investigate the self-similarity of feature networks constructed at different hidden layers, and analyzes how adjusting the degree of self-similarity in feature networks can enhance the classification performance of deep neural networks. Validated on three types of networks MLP architectures, convolutional networks, and attention architectures this study reveals that the degree of self-similarity exhibited by feature networks varies across different model architectures. Furthermore, embedding constraints on the self-similarity of feature networks during the training process can improve the performance of self-similar deep neural networks (MLP architectures and attention architectures) by up to 6 percentage points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。