提出非梯度的卷曲学习机制,解释生物神经网络如何不依赖梯度仍能高效优化。
Curl Descent: Non-Gradient Learning Dynamics with Sign-Diverse Plasticity
- 引入卷曲项模拟非梯度学习动力学,源自抑制-兴奋连接或赫布/反赫布可塑性。
- 小卷曲项保持稳定,大卷曲项导致混沌或意外加速学习,突破鞍点。
- 适用于研究生物学习机制或设计鲁棒性更强的非梯度学习算法的人。
梯度驱动算法是人工神经网络训练的核心,但生物神经网络在学习中是否采用类似策略尚不明确。实验常发现突触可塑性规则多样,却不清楚这些规则是否近似于梯度下降。本文探讨一种被忽视的可能性:学习动力学可能包含根本性的非梯度“卷曲”成分,同时仍能有效优化损失函数。具有抑制-兴奋连接或赫布/反赫布可塑性的网络自然产生卷曲项,使学习动力学无法被表述为任何目标的梯度下降。我们在可解析的师生框架下分析前馈网络,系统引入规则反转的神经元以施加非梯度动力学。小卷曲项维持原解流形的稳定性,学习行为类似梯度下降;超过临界值后,强卷曲项破坏稳定性,导致混沌学习并损害性能。但在某些网络架构中,卷曲项反而通过暂时上升损失,帮助权重逃离鞍点,实现比梯度下降更快的学习。结果揭示了支持多样化学习规则的稳健架构,为神经网络梯度学习的规范理论提供了重要反例。
原文摘要 · Abstract (English)
Gradient-based algorithms are a cornerstone of artificial neural network training, yet it remains unclear whether biological neural networks use similar gradient-based strategies during learning. Experiments often discover a diversity of synaptic plasticity rules, but whether these amount to an approximation to gradient descent is unclear. Here we investigate a previously overlooked possibility: that learning dynamics may include fundamentally non-gradient "curl"-like components while still being able to effectively optimize a loss function. Curl terms naturally emerge in networks with inhibitory-excitatory connectivity or Hebbian/anti-Hebbian plasticity, resulting in learning dynamics that cannot be framed as gradient descent on any objective. To investigate the impact of these curl terms, we analyze feedforward networks within an analytically tractable student-teacher framework, systematically introducing non-gradient dynamics through neurons exhibiting rule-flipped plasticity. Small curl terms preserve the stability of the original solution manifold, resulting in learning dynamics similar to gradient descent. Beyond a critical value, strong curl terms destabilize the solution manifold. Depending on the network architecture, this loss of stability can lead to chaotic learning dynamics that destroy performance. In other cases, the curl terms can counterintuitively speed learning compared to gradient descent by allowing the weight dynamics to escape saddles by temporarily ascending the loss. Our results identify specific architectures capable of supporting robust learning via diverse learning rules, providing an important counterpoint to normative theories of gradient-based learning in neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。