arXiv:2510.16851cs.CLcs.AI2025-10被引 1

通过神经元分组通信,让大模型更高效、可解释地学习复杂推理。

Neuronal Group Communication for Efficient Neural representation

  • 将神经网络视为神经元组的动态交互系统,用低秩信号传递替代独立权重训练。
  • 在大语言模型上实现中等压缩率下复杂推理任务性能提升,优于传统低秩方法。
  • 提出神经稳定性度量,揭示推理能力来自外部驱动但保持动态稳定。

现代神经网络规模持续增长,带来性能提升的同时也面临效率与可解释性挑战。本文提出神经元组通信(NGC)框架,将神经网络重新构想为相互作用的神经元组动态系统,而非孤立权重的集合。权重被视作嵌入态之间的瞬时交互,计算通过组间迭代通信完成。该低秩模块化表示大幅减少冗余参数:神经元组间交换低维信号,实现组内专精与组间信息共享。基于动力系统理论,引入类似李雅普诺夫稳定的神经元稳定性度量,量化序列处理中激活向稳定模式收缩的程度。实验发现,涌现的推理能力对应于外部驱动力或“势能”,推动动态偏离平凡轨迹但维持稳定性。在大语言模型上实证表明,NGC在中等压缩率下显著提升复杂推理表现,优于标准低秩近似和跨层基共享方法。最后讨论了结构化神经元组动态对高维学习系统泛化能力的潜在影响。

原文摘要 · Abstract (English)

The ever-increasing scale of modern neural networks has brought unprecedented performance alongside daunting challenges in efficiency and interpretability. This paper addresses the core question of how to build large neural systems that learn efficient, modular, and interpretable representations. We propose Neuronal Group Communication (NGC), a theory-driven framework that reimagines a neural network as a dynamical system of interacting neuronal groups rather than a monolithic collection of neural weights. Instead of treating each weight as an independent trainable parameter, NGC treats weights as transient interactions between embedding-like neuronal states, with neural computation unfolding through iterative communication among groups of neurons. This low-rank, modular representation yields compact models: groups of neurons exchange low-dimensional signals, enabling intra-group specialization and inter-group information sharing while dramatically reducing redundant parameters. By drawing on dynamical systems theory, we introduce a neuronal stability metric (analogous to Lyapunov stability) that quantifies the contraction of neuron activations toward stable patterns during sequence processing. Using this metric, we reveal that emergent reasoning capabilities correspond to an external driving force or ``potential'', which nudges the neural dynamics away from trivial trajectories while preserving stability. Empirically, we instantiate NGC in large language models (LLMs) and demonstrate improved performance on complex reasoning benchmarks under moderate compression. NGC consistently outperforms standard low-rank approximations and cross-layer basis-sharing methods at comparable compression rates. We conclude by discussing the broader implications of NGC, including how structured neuronal group dynamics might relate to generalization in high-dimensional learning systems.

神经网络大模型可解释性动态系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。