arXiv:2501.08341cond-mat.dis-nncond-mat.stat-mech2025-01

用最简神经网络解析学习机制与损失曲面特性

Dissecting a Small Artificial Neural Network

  • 分析九维参数空间中的损失曲面结构特征
  • 发现权重持续漂移但能快速收敛至零损失
  • 揭示学习如退火过程,具相变类似热力学现象

我们研究了表示逻辑异或(XOR)门的最简单人工神经网络的损失曲面和反向传播收敛动态。在九维参数空间中,损失曲面的截面呈现出显著特征,有助于理解为何反向传播能高效实现零损失收敛,而权重和偏置值却持续漂移。对比非随机与随机批次所得截面形状差异。借鉴统计物理,引入微正则熵作为表征网络相变行为的独特量。由此可知,神经网络学习可视为经历类比热力学相变的退火过程。同时揭示:随着隐藏神经元增多,损失曲面简化,消除由有限尺寸效应引发的熵垒。

原文摘要 · Abstract (English)

We investigate the loss landscape and backpropagation dynamics of convergence for the simplest possible artificial neural network representing the logical exclusive-OR (XOR) gate. Cross-sections of the loss landscape in the nine-dimensional parameter space are found to exhibit distinct features, which help understand why backpropagation efficiently achieves convergence toward zero loss, whereas values of weights and biases keep drifting. Differences in shapes of cross-sections obtained by nonrandomized and randomized batches are discussed. In reference to statistical physics we introduce the microcanonical entropy as a unique quantity that allows to characterize the phase behavior of the network. Learning in neural networks can thus be thought of as an annealing process that experiences the analogue of phase transitions known from thermodynamic systems. It also reveals how the loss landscape simplifies as more hidden neurons are added to the network, eliminating entropic barriers caused by finite-size effects.

神经网络损失曲面学习机制相变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。