提出Sign-In方法,通过动态重参数化解决稀疏网络从零训练的初始化难题。
Sign-In to the Lottery: Reparameterizing Sparse Training From Scratch
- 用动态重参数化实现参数符号的可证明翻转,提升稀疏训练初始质量。
- 实验与理论表明,该方法能显著改善从零开始训练稀疏网络的性能。
- 适合关注稀疏训练优化、模型压缩与高效深度学习的研究者。
从零训练稀疏神经网络(PaI)与密集到稀疏训练之间的性能差距,是高效深度学习的主要障碍。根据彩票券假说,PaI 的关键在于找到特定问题的参数初始化。我们发现,确定正确的参数符号已足够。然而,这些符号在 PaI 中仍难以获得。为此,我们提出 Sign-In,采用一种动态重参数化机制,可证明地诱导参数符号翻转。这种符号翻转与密集到稀疏训练所能实现的互补,使 Sign-In 成为一种正交方法。我们的实验和理论表明,Sign-In 能有效提升 PaI 性能,同时也揭示了缩小 PaI 与密集到稀疏训练差距的核心挑战。
原文摘要 · Abstract (English)
The performance gap between training sparse neural networks from scratch (PaI) and dense-to-sparse training presents a major roadblock for efficient deep learning. According to the Lottery Ticket Hypothesis, PaI hinges on finding a problem specific parameter initialization. As we show, to this end, determining correct parameter signs is sufficient. Yet, they remain elusive to PaI. To address this issue, we propose Sign-In, which employs a dynamic reparameterization that provably induces sign flips. Such sign flips are complementary to the ones that dense-to-sparse training can accomplish, rendering Sign-In as an orthogonal method. While our experiments and theory suggest performance improvements of PaI, they also carve out the main open challenge to close the gap between PaI and dense-to-sparse training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。