提出可微稀疏方法LaPrune,实现百万级模型的精确稀疏控制。
LaPrune: Controllable Differentiable Sparsity at Million Scale

- 通过拉普拉斯和约束,精确控制稀疏度与质量分布。
- 在固定预算下,使掩码趋于硬性Top-k选择,同时保持选中质量不变。
- 适用于大规模模型训练,对评分尺度不敏感,适合需要稳定稀疏的场景。
Top-$k$选择决定稀疏模型中哪些组件保持激活。硬性选择会阻断梯度,而连续松弛通常将掩码硬度与选中质量耦合。我们提出LaPrune,一种数学上精确控制预算的可微层,可在保持选中质量的同时控制归一化二阶矩。拉普拉斯和屏障确保选择质量不变,归一化二阶矩约束使掩码从密集均质分配向硬性Top-$k$移动。我们推导出饱和比例的群体预测、近二值极限律,以及近零分数的紧致最坏情况保证。归一化硬度参数对评分尺度不变,而固定拉普拉斯温度则不然。
原文摘要 · Abstract (English)
Top-$k$ selection determines which components of a sparse model remain active. Hard selection blocks gradients, while continuous relaxations often couple mask hardness to the selected mass. We introduce LaPrune, a mathematically exact-budget differentiable layer that controls the normalized second moment while preserving the selected mass. A LapSum barrier preserves the selection mass, and a normalized second-moment constraint moves the mask from a dense equal-mass allocation toward hard top-$k$ at each budget. We derive a population prediction of the saturated fraction, a near-binary limiting law, and a tight worst-case guarantee on the near-zero fraction. The normalized hardness parameter is invariant to score scale, while a fixed LapSum temperature is not.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。