arXiv:2501.02436cs.LGnlin.CD2025-01

用动力系统理论解析深度学习训练过程中的相变现象

Network Dynamics-Based Framework for Understanding Deep Neural Networks

  • 从神经元级定义保序与非保序变换,揭示网络动态行为机制
  • 发现训练中存在可解释的分相转变,对应如'领悟'等关键现象
  • 适合研究模型泛化、架构优化与训练策略的学者参考

人工智能的发展亟需深入理解深度学习的本质机制。本文提出一个基于动力系统理论的分析框架,通过引入神经元层面的两类基本变换——保序变换与非保序变换,重新定义神经网络中的线性与非线性。不同变换模式导致权重向量组织方式、信息提取模式及学习阶段的显著差异,训练过程中这些阶段之间的跃迁可解释诸如'领悟'(grokking)等关键现象。为进一步刻画泛化能力与结构稳定性,本文引入样本空间与权重空间中的吸引子盆地概念。各层神经元中不同类型变换的分布,以及两类吸引子盆地的结构特征,构成一组核心性能分析指标。深度、宽度、学习率、批大小等超参数则作为调控这些指标的控制变量。该框架不仅揭示了深度学习的内在优势,也为网络架构与训练策略优化提供了新视角。

原文摘要 · Abstract (English)

Advancements in artificial intelligence call for a deeper understanding of the fundamental mechanisms underlying deep learning. In this work, we propose a theoretical framework to analyze learning dynamics through the lens of dynamical systems theory. We redefine the notions of linearity and nonlinearity in neural networks by introducing two fundamental transformation units at the neuron level: order-preserving transformations and non-order-preserving transformations. Different transformation modes lead to distinct collective behaviors in weight vector organization, different modes of information extraction, and the emergence of qualitatively different learning phases. Transitions between these phases may occur during training, accounting for key phenomena such as grokking. To further characterize generalization and structural stability, we introduce the concept of attraction basins in both sample and weight spaces. The distribution of neurons with different transformation modes across layers, along with the structural characteristics of the two types of attraction basins, forms a set of core metrics for analyzing the performance of learning models. Hyperparameters such as depth, width, learning rate, and batch size act as control variables for fine-tuning these metrics. Our framework not only sheds light on the intrinsic advantages of deep learning, but also provides a novel perspective for optimizing network architectures and training strategies.

深度学习动力系统学习相变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。