用凹函数正则化找稀疏神经网络中的高效子结构
Playing the Lottery With Concave Regularizers for Sparse Trainable Neural Networks
- 用凹正则化引导网络拓扑的稀疏化,提升稀疏子网搜索效率
- 在多个数据集和架构上优于现有方法,能发现性能更优的稀疏子网
- 适合关注模型压缩与高效训练的研究者或工程师
稀疏神经网络因其在推理阶段显著降低计算与存储开销而受到广泛关注。在此背景下,彩票理论(Lottery Ticket Hypothesis, LTH)指出:存在可独立训练且表现优异的稀疏子网络,称为中奖票(winning tickets)。如何高效寻找这些中奖票仍是开放问题。本文提出一类新方法,核心是使用凹正则化来促进松弛二值掩码(表示网络拓扑)的稀疏性。我们在凸框架下对方法有效性进行了理论分析,并在多种数据集与网络架构上进行了扩展数值实验,结果表明该方法能有效提升现有先进算法的性能。
原文摘要 · Abstract (English)
The design of sparse neural networks, i.e., of networks with a reduced number of parameters, has been attracting increasing research attention in the last few years. The use of sparse models may significantly reduce the computational and storage footprint in the inference phase. In this context, the lottery ticket hypothesis (LTH) constitutes a breakthrough result, that addresses not only the performance of the inference phase, but also of the training phase. It states that it is possible to extract effective sparse subnetworks, called winning tickets, that can be trained in isolation. The development of effective methods to play the lottery, i.e., to find winning tickets, is still an open problem. In this article, we propose a novel class of methods to play the lottery. The key point is the use of concave regularization to promote the sparsity of a relaxed binary mask, which represents the network topology. We theoretically analyze the effectiveness of the proposed method in the convex framework. Then, we propose extended numerical tests on various datasets and architectures, that show that the proposed method can improve the performance of state-of-the-art algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。