arXiv:2509.23898cs.LGstat.ML2025-09NeurIPS被引 3

提出可微分的结构化稀疏方法,让模型自动学习压缩且无需额外剪枝。

Differentiable Sparsity via $D$-Gating: Simple and Versatile Structured Penalization

  • 用门控因子分解权重,实现可微的结构化稀疏正则化
  • 理论证明最小值等价于传统稀疏目标,收敛速度指数级提升
  • 在视觉、语言、表格任务中均优于传统剪枝和直接优化方法

结构化稀疏正则化能有效压缩神经网络,但其不可微性破坏了与标准随机梯度下降的兼容性,需依赖特殊优化器或事后剪枝,缺乏理论保证。本文提出 $D$-Gating,一种完全可微的结构化过参数化方法,将每组权重拆分为主权重向量和多个标量门控因子。我们证明,在 $D$-Gating 下的任意局部最小值,也对应于非光滑结构化 $L_{2,2/D}$ 正则化的目标的局部最小值,并进一步表明 $D$-Gating 目标在梯度流极限下至少以指数速度收敛到 $L_{2,2/D}$ 正则化损失。综合结果表明,$D$-Gating 在理论上等价于求解原始组稀疏问题,同时诱导出从非稀疏到稀疏优化的独特学习动态。我们在视觉、语言和表格任务上验证了该理论,结果显示 $D$-Gating 始终在性能与稀疏性之间取得优异平衡,显著优于直接优化结构化惩罚和传统剪枝基线。

原文摘要 · Abstract (English)

Structured sparsity regularization offers a principled way to compact neural networks, but its non-differentiability breaks compatibility with conventional stochastic gradient descent and requires either specialized optimizers or additional post-hoc pruning without formal guarantees. In this work, we propose $D$-Gating, a fully differentiable structured overparameterization that splits each group of weights into a primary weight vector and multiple scalar gating factors. We prove that any local minimum under $D$-Gating is also a local minimum using non-smooth structured $L_{2,2/D}$ penalization, and further show that the $D$-Gating objective converges at least exponentially fast to the $L_{2,2/D}$-regularized loss in the gradient flow limit. Together, our results show that $D$-Gating is theoretically equivalent to solving the original group sparsity problem, yet induces distinct learning dynamics that evolve from a non-sparse regime into sparse optimization. We validate our theory across vision, language, and tabular tasks, where $D$-Gating consistently delivers strong performance-sparsity tradeoffs and outperforms both direct optimization of structured penalties and conventional pruning baselines.

稀疏化可微分结构化模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。