强化学习让神经网络自发形成复杂动态计算机制。
Emergence of hybrid computational dynamics through reinforcement learning
- 用强化学习训练循环网络,自发产生混合吸引子结构。
- 相比监督学习,强化学习的解更稳定且性能更好,尤其在复杂任务中。
- 适合研究智能系统动态机制或设计自适应AI的科研人员。
理解学习算法如何塑造神经网络中的计算策略,是机器智能的核心挑战。尽管网络架构备受关注,但学习范式本身对涌现动态的影响仍不明确。本文发现,在相同决策任务下,强化学习(RL)与监督学习(SL)使循环神经网络(RNNs)走向根本不同的计算解。通过系统的动力学分析,我们揭示RL自发形成混合吸引子架构:稳定固定点用于维持决策,准周期吸引子用于灵活整合证据。这与SL几乎仅收敛于简单固定点解形成鲜明对比。此外,RL通过一种强大的隐式正则化,塑造功能平衡的神经种群,增强鲁棒性,而SL解则表现出更高异质性。这些复杂动态在RL中的出现程度受权重初始化调控,并与性能提升显著相关,尤其在任务复杂度增加时。结果表明,学习算法是决定涌现计算的关键因素,奖励驱动优化能自主发现梯度优化难以触及的复杂动力学机制。这些发现为神经计算提供了机理洞察,并为设计自适应人工智能系统提供可操作原则。
原文摘要 · Abstract (English)
Understanding how learning algorithms shape the computational strategies that emerge in neural networks remains a fundamental challenge in machine intelligence. While network architectures receive extensive attention, the role of the learning paradigm itself in determining emergent dynamics remains largely unexplored. Here we demonstrate that reinforcement learning (RL) and supervised learning (SL) drive recurrent neural networks (RNNs) toward fundamentally different computational solutions when trained on identical decision-making tasks. Through systematic dynamical systems analysis, we reveal that RL spontaneously discovers hybrid attractor architectures, combining stable fixed-point attractors for decision maintenance with quasi-periodic attractors for flexible evidence integration. This contrasts sharply with SL, which converges almost exclusively to simpler fixed-point-only solutions. We further show that RL sculpts functionally balanced neural populations through a powerful form of implicit regularization -- a structural signature that enhances robustness and is conspicuously absent in the more heterogeneous solutions found by SL-trained networks. The prevalence of these complex dynamics in RL is controllably modulated by weight initialization and correlates strongly with performance gains, particularly as task complexity increases. Our results establish the learning algorithm as a primary determinant of emergent computation, revealing how reward-based optimization autonomously discovers sophisticated dynamical mechanisms that are less accessible to direct gradient-based optimization. These findings provide both mechanistic insights into neural computation and actionable principles for designing adaptive AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。