arXiv:2602.23405cs.NEcs.LG2026-02

用各向同性激活函数实现神经网络的动态重构与稀疏化。

Isotropic Activation Functions Enable Deindividuated Neurons and Adaptive Topologies

  • 通过对称性重参数化与奇异值分解,将层变为一一对应的有序连接。
  • 实现渐近50%参数稀疏化,功能保持不变,支持实时结构调整。
  • 适合需要可解释性与自适应架构的深度学习应用研究者。

本文提出一种基于各向同性激活函数的密集神经网络拓扑自适应方法。通过预设的重参数化对称性和仿射映射的奇异值分解,将网络层对角化为一一对应、有序的连接,简化了单个连接对函数影响的评估。低影响神经元可被移除(神经退化),同时保留阈值缓冲区中大量不活跃的“骨架”神经元(神经发生)。这些由对称性引导的对角化与结构变化具有函数不变性:神经发生时计算等价,神经退化时可任意逼近,实现密集网络在功能不变前提下的渐近50%参数稀疏化。因此,可实现实时响应任务需求、任务增删或变更的架构重构。该方法以原始对称性约束为核心,导出具备显式基无关性的各向同性函数,隐含神经元个体化损失,使层可自由在不同基下分解与解释,直接支持自适应拓扑。此外,引入新可调参数“内在长度”以增强分析不变性,并提出广义各向同性感知机架构,支持所有矩阵-向量积的并行预计算,呈现嵌套函数类。对角化为各向同性网络的可解释性与监控提供了新可能。

原文摘要 · Abstract (English)

Introduced is a methodology for adapting the topology of dense neural networks, enabled by isotropic activation functions. Achieved through prescribed reparameterisation symmetries and singular-value decomposition of affine maps, this diagonalises layers into one-to-one, ordered connections. This makes it simpler to assess the impact of individual connections on the function. Low-impact neurons can be removed (neurodegeneration), and a thresholded buffer of largely inactive 'scaffold' neurons is maintained (neurogenesis). These symmetry-led diagonalisation and structural changes are function-invariant, demonstrated to be computationally identical during neurogenesis, arbitrarily well approximated during neurodegeneration, and enable asymptotic 50\% parameter sparsification of dense networks with identically preserved function. Thus, real-time restructuring of the architecture in response to task demands, task appending, removal or changes is shown. The approach is conceptually centred on primitive symmetry-prescriptions, through which isotropic functions are derived that feature explicit basis independence and a loss in the individuation of neurons implicit in typical elementwise functional forms. Hence, this allows freedom in the basis to which layers are decomposed and interpreted as individual artificial neurons, directly enabling this adaptive topology approach. Additionally, a new tunable model parameter, the 'intrinsic length', is introduced to improve this analytical invariance, alongside a generalised isotropic-perceptron architecture that enables parallel precomputation of all matrix-vector products and displays a nested functional class. Diagonalisation is suggested to offer new possibilities for interpretability and monitoring of isotropic networks.

神经网络自适应架构稀疏化可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。