arXiv:2501.09659cs.LG2025-01

用统计物理方法模拟神经网络权重演化,揭示训练背后的数学规律。

Fokker-Planck to Callan-Symanzik: evolution of weight matrices under training

  • 基于福克-普朗克方程建模权重矩阵概率分布演化
  • 在双瓶颈自编码器中验证理论与实测分布一致性
  • 推导出卡兰-西曼兹克等物理方程,连接深度学习与场论

神经网络训练过程的动力学演化是极具吸引力的研究课题。将统计物理中变量演化的首项原理应用于训练动态,虽概念上有效,但实际需数值求解如福克-普朗克方程,而全网模拟常遭遇维度灾难。本文利用福克-普朗克方程,模拟一个双瓶颈层自编码器中瓶颈层权重矩阵的概率密度演化,并通过检查输出数据分布对比理论与实测结果。此外,从所导出的动力学方程中,还推导出卡兰-西曼兹克、卡达尔-帕里西-宗方等物理相关偏微分方程。

原文摘要 · Abstract (English)

The dynamical evolution of a neural network during training has been an incredibly fascinating subject of study. First principal derivation of generic evolution of variables in statistical physics systems has proved useful when used to describe training dynamics conceptually, which in practice means numerically solving equations such as Fokker-Planck equation. Simulating entire networks inevitably runs into the curse of dimensionality. In this paper, we utilize Fokker-Planck to simulate the probability density evolution of individual weight matrices in the bottleneck layers of a simple 2-bottleneck-layered auto-encoder and compare the theoretical evolutions against the empirical ones by examining the output data distributions. We also derive physically relevant partial differential equations such as Callan-Symanzik and Kardar-Parisi-Zhang equations from the dynamical equation we have.

神经网络动力学统计物理权重演化偏微分方程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。