通过非均匀连接结构提升稀疏网络性能,发现枢纽神经元位置比连接密度差异更重要。
Heterogeneous Connectivity in Sparse Networks: Fan-in Profiles, Gradient Hierarchy, and Topological Equilibria

- 用连续非线性函数定义神经元连接偏好,实现稠密与稀疏感受野共存
- 90%稀疏度下准确率仅比稠密模型低0.2-0.6%,且结构化连接无优势
- 优化驱动的枢纽定位优于随机分配,动态训练中初始分布越接近平衡越高效
Profiled Sparse Networks (PSN) 以确定性的非均匀入度分布替代传统均匀连接,通过连续非线性函数定义神经元连接模式,使单个神经元同时具备稠密和稀疏的感受野。我们在四个分类数据集(涵盖视觉与表格领域,输入维度54至784,网络深度2-3层)上评估了PSN性能。在90%稀疏度下,所有静态连接模式(包括均匀随机基线)在各数据集上的准确率均仅比稠密基线低0.2-0.6%,表明当枢纽位置未与任务对齐时,非均匀连接并无精度优势。该结论在80%-99.9%稀疏度、八类参数化分布(含对数正态与幂律)、入度变异系数0至2.5范围内均成立。内部梯度分析显示,结构化分布使枢纽神经元梯度集中度达2-5倍,而随机基线仅为约1倍,层级强度与入度变异系数相关(r=0.93)。将PSN入度分布用于RigL动态稀疏训练初始化,对数正态分布匹配平衡入度分布时表现最优,优势随任务难度增加:在Fashion-MNIST上提升+0.16%(p=0.036, d=1.07),EMNIST上+0.43%,Forest Cover上+0.49%。RigL收敛至特征入度分布,无论初始状态如何。从平衡点开始可使优化器聚焦权重精调而非拓扑重排。枢纽神经元的选择比连接方差更重要——随机定位无优势,而优化驱动定位有效。
原文摘要 · Abstract (English)
Profiled Sparse Networks (PSN) replace uniform connectivity with deterministic, heterogeneous fan-in profiles defined by continuous, nonlinear functions, creating neurons with both dense and sparse receptive fields. We benchmark PSN across four classification datasets spanning vision and tabular domains, input dimensions from 54 to 784, and network depths of 2--3 hidden layers. At 90% sparsity, all static profiles, including the uniform random baseline, achieve accuracy within 0.2-0.6% of dense baselines on every dataset, demonstrating that heterogeneous connectivity provides no accuracy advantage when hub placement is arbitrary rather than task-aligned. This result holds across sparsity levels (80-99.9%), profile shapes (eight parametric families, lognormal, and power-law), and fan-in coefficients of variation from 0 to 2.5. Internal gradient analysis reveals that structured profiles create a 2-5x gradient concentration at hub neurons compared to the ~1x uniform distribution in random baselines, with the hierarchy strength predicted by fan-in coefficient of variation ($r = 0.93$). When PSN fan-in distributions are used to initialise RigL dynamic sparse training, lognormal profiles matched to the equilibrium fan-in distribution consistently outperform standard ERK initialisation, with advantages growing on harder tasks, achieving +0.16% on Fashion-MNIST ($p = 0.036$, $d = 1.07$), +0.43% on EMNIST, and +0.49% on Forest Cover. RigL converges to a characteristic fan-in distribution regardless of initialisation. Starting at this equilibrium allows the optimiser to refine weights rather than rearrange topology. Which neurons become hubs matters more than the degree of connectivity variance, i.e., random hub placement provides no advantage, while optimisation-driven placement does.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。