强制神经网络学习低维表示,可提升时间序列预测的泛化能力。
Emergent Generalization by Representation Learning in Artificial Neural Networks

- 通过信息瓶颈迫使网络学习低维表示,实现旋转与分布外泛化。
- 低维表示的涌现动态呈非单调变化,峰值与泛化性能强相关。
- 实验发现小鼠海马区活动也存在类似动态,支持其认知功能意义。
降维在识别神经流形方面表现强大,这些低维结构隐藏于高维神经活动中,提升了群体编码的可解释性。然而,这类低维表示是否具有生物学意义并带来学习系统的功能优势,抑或仅反映神经元层面活动,仍存争议。我们发现,在时间序列预测任务中,显式的信息瓶颈迫使循环神经网络学习低维表示,是实现旋转泛化与分布外泛化所必需的。利用因果涌现的信息论度量,我们刻画了该表示在记忆到泛化的转变过程中的动态,发现其轨迹呈非单调变化:先下降、达最小值后上升至峰值,而预测损失则持续下降。该轨迹随任务复杂度增加而扩展,涌现结构的强度可稳定预测泛化性能。对小鼠学习交替迷宫任务时海马CA1区活动的分析揭示了类似的非单调涌现动态,且与行为表现同步。这些结果表明,神经网络学习紧凑、分布式且涌现的表示具有功能性优势,支持其在认知中的因果作用。
原文摘要 · Abstract (English)
Dimensionality reduction has proven powerful for identifying neural manifolds, which are low-dimensional structures underlying high-dimensional neural activity. These low-dimensional representations have improved the interpretability of population-level coding. Yet whether such low-dimensional representations are biologically relevant and confer functional advantages in learning systems, or merely reflect neuron-level activity, remains contested in neuroscience. We show that an explicit information bottleneck forcing a recurrent neural network to learn a low-dimensional representation is necessary for rotational and out-of-distribution generalisation in a time-series prediction task. Using information-theoretic measures of causal emergence, we characterise the dynamics of this representation across the memorisation-to-generalisation transition, finding a non-monotonic trajectory which shows an initial decrease, a minimum, and a subsequent rise to a maximum, even as prediction loss falls monotonically. This trajectory scales with task complexity, and the magnitude of emergent structure reliably predicts generalisation performance. Analysis of CA1 hippocampal activity in mice learning an alternating maze task reveals analogous non-monotonic emergence dynamics that track behavioural performance. Together, these findings indicate that the ability of neural networks to learn compact, distributed and emergent representations confers a functional advantage for generalisation, supporting a causal role for learned representations in cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。