提出新型神经网络,让记忆更稳定连续。
Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory

- 用突触除法归一化构建动态网络,避免状态破碎。
- 训练后有效秩显著下降,收敛到低维慢流形。
- 适合研究连续记忆的神经机制或模型优化者。
维持和更新连续变量是工作记忆的核心能力。经典连续吸引子网络对微调极度敏感,而标准递归神经网络(如GRU、LSTM)通常无法稳定学习连续流形,反而将状态空间分裂为离散点吸引子。为此,我们借鉴皮层回路中广泛存在的除法归一化机制,提出递归除法归一化网络(RDNN),一种极简且代数隔离的动态除法模型。通过对典型工作记忆任务的动力系统分析,我们证明该生物物理约束使网络能收敛至鲁棒、高保真的慢流形。进一步分析反向传播时的梯度动态,发现除法归一化引入了活动依赖的局部梯度缩放,抑制高激活区域的参数更新。这与网络有效秩显著自压缩的现象一致,使循环动态被限制在紧密的低维子空间内,避免了显式低秩分解相关的优化病态。消融实验表明,虽减性抑制可维持静态记忆,但只有除法归一化在时变输入下能数学上防止流形破碎。研究揭示除法归一化不仅是生物现象,更是学习高保真连续表征的关键计算机制。
原文摘要 · Abstract (English)
The ability to robustly maintain and update continuous variables is a hallmark of working memory. While classical continuous attractor networks suffer from severe fine-tuning fragility, standard artificial recurrent neural networks (RNNs) like GRUs and LSTMs typically fail to stably learn continuous manifolds, instead shattering the state space into discretized point attractors. To bridge this gap, we draw inspiration from divisive normalization, a canonical neural computation widely observed across cortical circuits, and propose the Recurrent Divisive Normalization Network (RDNN), a minimal and algebraically isolated model of dynamic division. Through dynamical systems analysis on canonical working memory tasks, we demonstrate that this biophysical constraint allows the network to converge to robust, high-fidelity slow manifolds. Furthermore, we analyze the gradient dynamics of divisive normalization during Backpropagation Through Time (BPTT), showing that it introduces an activity-dependent local gradient scaling. This scaling dampens parameter updates in highly active regimes, which empirically aligns with a significant self-compression of the network's effective rank, confining the recurrent dynamics to a tight, low-dimensional subspace while avoiding the optimization pathologies associated with explicit low-rank factorization. Finally, ablations demonstrate that while subtractive inhibition can maintain static memories, divisive normalization is mathematically essential to prevent manifold shattering under time-varying inputs. Our findings identify divisive normalization not merely as a biological artifact, but as a critical computational mechanism for learning high-fidelity continuous representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。