arXiv:2409.03749cs.LGq-bio.NC2024-09NeurIPS

研究非线性感知机中监督与强化学习的动态机制。

Dynamics of Supervised and Reinforcement Learning in the Non-Linear Perceptron

  • 用随机过程推导非线性感知机的学习流方程。
  • 发现输入噪声对监督/强化学习速度影响不同。
  • 适用于复杂神经网络架构的动态分析方法。

神经网络的学习效率取决于任务结构和学习规则。以往研究多基于线性输出或师生框架下的感知机,虽利于理论分析,却难以揭示非线性与输入分布对学习动态的真实影响。本文采用随机过程方法,推导非线性感知机进行二分类时的学习流方程。结果表明,学习规则(监督学习/强化学习)与输入数据分布共同决定学习曲线和遗忘曲线。特别是,输入噪声对监督学习与强化学习的学习速度产生不同影响,并决定任务间学习的覆盖速度。我们通过MNIST数据集验证了该方法的有效性。该框架为分析更复杂电路架构的学习动态提供了新路径。

原文摘要 · Abstract (English)

The ability of a brain or a neural network to efficiently learn depends crucially on both the task structure and the learning rule. Previous works have analyzed the dynamical equations describing learning in the relatively simplified context of the perceptron under assumptions of a student-teacher framework or a linearized output. While these assumptions have facilitated theoretical understanding, they have precluded a detailed understanding of the roles of the nonlinearity and input-data distribution in determining the learning dynamics, limiting the applicability of the theories to real biological or artificial neural networks. Here, we use a stochastic-process approach to derive flow equations describing learning, applying this framework to the case of a nonlinear perceptron performing binary classification. We characterize the effects of the learning rule (supervised or reinforcement learning, SL/RL) and input-data distribution on the perceptron's learning curve and the forgetting curve as subsequent tasks are learned. In particular, we find that the input-data noise differently affects the learning speed under SL vs. RL, as well as determines how quickly learning of a task is overwritten by subsequent learning. Additionally, we verify our approach with real data using the MNIST dataset. This approach points a way toward analyzing learning dynamics for more-complex circuit architectures.

学习动态非线性感知机监督学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。