通过生物启发机制调控神经元连接,加速模型从记忆到泛化的转变。
Biologically Inspired Mechanisms for Facilitating Grokking in Multilayer Perceptrons

- 引入输入门控、结构可塑性等七种生物机制调节隐藏层计算。
- 稳态机制显著提升泛化能力,结构稀疏化次之,其他机制效果较弱。
- 适用于大模型训练优化,或可缩短泛化所需训练时间。
Grokking 是一种延迟的从记忆到泛化的转变,常伴随内部表征的显著重组。本文研究了多种未被普遍纳入人工神经网络的生物启发机制,是否能通过调节神经元活动、响应及有效连接,主动促进这一转变。我们为多层感知机引入输入门控、结构可塑性、增益调制、阈值调制、稳态、侧抑制和激活去相关,并在两个经典 grokking 基准任务——稀疏奇偶校验和噪声 XOR 分类上进行系统消融实验。结果表明,这些机制对泛化的贡献不均:稳态提供最强且最一致的增益,结构稀疏化是第二关键机制;其余机制在当前实验中作用较小或不够稳定。两类问题的结果均支持一个共同原则:显式调控神经元使用率与有效连接,可改善可泛化的内部计算涌现。这些发现推动对生物启发的活动调控与自适应稀疏化的更广泛探索,尤其在大语言模型中,可能加速通用表征的形成并减少实现稳健泛化的优化时间。
原文摘要 · Abstract (English)
Grokking is a delayed transition from memorization to generalization that is often accompanied by substantial reorganization of internal representations. This paper studies whether biologically inspired mechanisms, many of which are not commonly incorporated into artificial neural networks, can actively promote this transition by regulating hidden-layer computation at the levels of neuronal activity, response, and effective connectivity. We augment a multilayer perceptron with input gating, structural plasticity, gain modulation, threshold modulation, homeostasis, lateral inhibition, and activation decorrelation, and evaluate these mechanisms through systematic ablations on two established grokking benchmarks: sparse parity and noisy XOR classification. The results show that the mechanisms contribute unequally to generalization. Homeostasis provides the strongest and most consistent benefit, while structural sparsification emerges as the second major mechanism. The remaining biologically inspired mechanisms have smaller or less consistent effects in the present experiments. For both problems, the results support the common principle that explicit regulation of neuron utilization and effective connectivity can improve the emergence of generalizable internal computation. These findings motivate broader investigation of biologically inspired activity regulation and adaptive sparsification, including in large language models, where they may accelerate the development of generalizable representations and reduce the optimization time required for robust generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。