arXiv:2502.06300cs.LG2025-02ICLR

研究如何分配有限可学习权重以提升神经网络表达能力

The impact of allocation strategies in subset learning on the expressive power of neural networks

  • 通过教师-学生框架量化不同权重分配的表达能力
  • 发现权重分布越广,网络表达力越强,尤其在线性网络中
  • 结论对深层ReLU网络也适用,适合模型优化研究者

传统机器学习中,模型由一组参数定义并针对特定任务进行优化。在神经网络中,这些参数对应突触权重。然而现实中往往无法控制或更新所有权重。这一挑战不仅存在于人工网络,也存在于生物网络(如大脑),其中学习过程中突触权重的分布式修改程度尚不明确。受此启发,我们从理论角度研究:在固定数量可学习权重的前提下,不同分配策略如何影响神经网络的容量。采用教师-学生设置,提出基准方法来量化每种分配对应的表达能力。我们确立了在线性循环神经网络和线性多层前馈网络中,使表达能力达到最大或最小的分配条件。对于次优分配,提出了启发式原则来估计其表达能力,该原则亦可推广至浅层ReLU网络。最后,通过实证实验验证了理论发现。结果强调了战略性地将可学习权重分布于网络中的关键作用,表明更广泛的分配通常能增强网络的表达能力。

原文摘要 · Abstract (English)

In traditional machine learning, models are defined by a set of parameters, which are optimized to perform specific tasks. In neural networks, these parameters correspond to the synaptic weights. However, in reality, it is often infeasible to control or update all weights. This challenge is not limited to artificial networks but extends to biological networks, such as the brain, where the extent of distributed synaptic weight modification during learning remains unclear. Motivated by these insights, we theoretically investigate how different allocations of a fixed number of learnable weights influence the capacity of neural networks. Using a teacher-student setup, we introduce a benchmark to quantify the expressivity associated with each allocation. We establish conditions under which allocations have maximal or minimal expressive power in linear recurrent neural networks and linear multi-layer feedforward networks. For suboptimal allocations, we propose heuristic principles to estimate their expressivity. These principles extend to shallow ReLU networks as well. Finally, we validate our theoretical findings with empirical experiments. Our results emphasize the critical role of strategically distributing learnable weights across the network, showing that a more widespread allocation generally enhances the network's expressive power.

神经网络表达能力权重分配理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。