arXiv:2608.04057cs.LGcs.AI2026-08

提出可微稀疏方法LaPrune,实现百万级模型的精确稀疏控制。

LaPrune: Controllable Differentiable Sparsity at Million Scale

论文配图:LaPrune: Controllable Differentiable Sparsity at Million Scale
图 1 · 摘自论文原文
  • 通过拉普拉斯和约束,精确控制稀疏度与质量分布。
  • 在固定预算下,使掩码趋于硬性Top-k选择,同时保持选中质量不变。
  • 适用于大规模模型训练,对评分尺度不敏感,适合需要稳定稀疏的场景。

Top-$k$选择决定稀疏模型中哪些组件保持激活。硬性选择会阻断梯度,而连续松弛通常将掩码硬度与选中质量耦合。我们提出LaPrune,一种数学上精确控制预算的可微层,可在保持选中质量的同时控制归一化二阶矩。拉普拉斯和屏障确保选择质量不变,归一化二阶矩约束使掩码从密集均质分配向硬性Top-$k$移动。我们推导出饱和比例的群体预测、近二值极限律,以及近零分数的紧致最坏情况保证。归一化硬度参数对评分尺度不变,而固定拉普拉斯温度则不然。

原文摘要 · Abstract (English)

Top-$k$ selection determines which components of a sparse model remain active. Hard selection blocks gradients, while continuous relaxations often couple mask hardness to the selected mass. We introduce LaPrune, a mathematically exact-budget differentiable layer that controls the normalized second moment while preserving the selected mass. A LapSum barrier preserves the selection mass, and a normalized second-moment constraint moves the mask from a dense equal-mass allocation toward hard top-$k$ at each budget. We derive a population prediction of the saturated fraction, a near-binary limiting law, and a tight worst-case guarantee on the near-zero fraction. The normalized hardness parameter is invariant to score scale, while a fixed LapSum temperature is not.

稀疏模型可微稀疏大规模训练优化算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。