arXiv:2512.06288cs.LGmath.ST2025-12被引 1

证明宽神经网络可压缩,为剪枝量化提供理论依据。

Theoretical Compression Bounds for Wide Multilayer Perceptrons

  • 提出随机贪心压缩算法,实现训练后剪枝与量化。
  • 证明宽MLP存在性能接近原模型的压缩子网。
  • 适用于宽MLP与CNN,无需数据假设,适合研究压缩理论者。

剪枝与量化技术在减少大型神经网络参数方面成效显著,但其经验成功缺乏理论支持。本文针对训练后剪枝与量化,提出一种随机贪心压缩算法,严格证明了多层感知机(MLP)存在性能优异的压缩子网。进一步将结果扩展至结构化剪枝及卷积神经网络(CNN),为宽网络剪枝提供了统一分析框架。理论结果不依赖数据假设,揭示了可压缩性与网络宽度间的权衡关系。所提算法与最优大脑损伤(OBD)有一定相似性,可视为其训练后随机版本。本研究弥合了剪枝/量化理论与应用之间的鸿沟,为宽多层感知机的压缩经验成功提供了理论解释。

原文摘要 · Abstract (English)

Pruning and quantization techniques have been broadly successful in reducing the number of parameters needed for large neural networks, yet theoretical justification for their empirical success falls short. We consider a randomized greedy compression algorithm for pruning and quantization post-training and use it to rigorously show the existence of pruned/quantized subnetworks of multilayer perceptrons (MLPs) with competitive performance. We further extend our results to structured pruning of MLPs and convolutional neural networks (CNNs), thus providing a unified analysis of pruning in wide networks. Our results are free of data assumptions, and showcase a tradeoff between compressibility and network width. The algorithm we consider bears some similarities with Optimal Brain Damage (OBD) and can be viewed as a post-training randomized version of it. The theoretical results we derive bridge the gap between theory and application for pruning/quantization, and provide a justification for the empirical success of compression in wide multilayer perceptrons.

神经网络压缩理论分析剪枝量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。