发现深度网络与大脑在学习时共享相同的物理规律。
Toward a Physics of Deep Learning and Brains
- 用非平衡统计物理描述网络活动级联,揭示学习最佳状态。
- 网络在准临界态下表现最优,而非严格临界点。
- 适用于优化模型性能,适合神经科学与AI交叉研究者。
深度神经网络与生物大脑都通过可调权重模拟突触连接,表面相似。本文发现,描述活体大脑神经爆发的方程同样适用于深度网络中的活动级联。这些方程源自非平衡统计物理,表明深度网络在吸收相与活跃相之间的临界区域学习效果最佳。由于强输入驱动,网络不处于真正临界点,而是在近似满足裂纹噪声标度关系的准临界态中运行。通过不同初始化训练,我们发现最大敏感性比接近临界点本身更能预测学习性能。结合有限尺寸标度分析,识别出包括巴克豪森噪声和定向渗流在内的不同普适类。该理论框架证明生物与人工神经网络共享普遍特征。
原文摘要 · Abstract (English)
Deep neural networks and brains both learn and share superficial similarities: processing nodes are likened to neurons and adjustable weights are likened to modifiable synapses. But can a unified theoretical framework be found to underlie them both? Here we show that the equations used to describe neuronal avalanches in living brains can also be applied to cascades of activity in deep neural networks. These equations are derived from non-equilibrium statistical physics and show that deep neural networks learn best when poised between absorbing and active phases. Because these networks are strongly driven by inputs, however, they do not operate at a true critical point but within a quasi-critical regime -- one that still approximately satisfies crackling noise scaling relations. By training networks with different initializations, we show that maximal susceptibility is a more reliable predictor of learning than proximity to the critical point itself. This provides a blueprint for engineering improved network performance. Finally, using finite-size scaling we identify distinct universality classes, including Barkhausen noise and directed percolation. This theoretical framework demonstrates that universal features are shared by both biological and artificial neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。