提出稀疏知识蒸馏的数学框架,解释为何分步压缩更有效。
Sparse Knowledge Distillation: A Mathematical Framework for Probability-Domain Temperature Scaling and Multi-Stage Compression
- 基于概率域软化算子构建统一理论框架
- 证明多阶段压缩收敛速度达 O(1/n),优于一次性剪枝
- 支持黑盒教师、部分访问等隐私保护场景
我们基于概率域软化算子,建立了一个统一的稀疏知识蒸馏理论框架。虽然 $p^{1/T} \propto \mathrm{softmax}(z/T)$ 的等价性广为人知,但本文贡献在于在此基础上构建算子级分析框架,而非等价本身。框架包含四个核心部分:(i)与算子无关的偏差-方差分解,揭示稀疏学生何时优于稠密教师;(ii)函数空间中多阶段剪枝的同伦路径形式化,解释为何迭代压缩成功而单次剪枝失败;(iii)收敛性保证,证明 n 阶蒸馏收敛率为 O(1/n),且参数依赖明确;(iv)等价类刻画,在容量约束下识别出产生相同学生模型的不同概率域算子。我们提出了基于排序保持、连续性、熵单调性、恒等性和边界行为的公理化定义,并证明多个非等价算子族满足该公理。所有学习理论保证均在该算子类上一致成立,与实现细节无关。这些结果为黑盒教师蒸馏、如 top-k 截断和仅文本输出等部分访问设置,以及隐私保护模型压缩提供了理论支撑。
原文摘要 · Abstract (English)
We develop a unified theoretical framework for sparse knowledge distillation based on probability-domain softening operators. While the equivalence $p^{1/T} \propto \mathrm{softmax}(z/T)$ is well known, our contribution is an operator-level analytical framework built on this foundation rather than the equivalence itself. The framework comprises four core components: (i) operator-agnostic bias--variance decompositions that characterize when sparse students outperform dense teachers, (ii) a homotopy path formalization of multi-stage pruning in function space explaining why iterative compression succeeds where one-shot pruning fails, (iii) convergence guarantees establishing $O(1/n)$ rates for $n$-stage distillation with explicit parameter dependence, and (iv) equivalence class characterizations identifying distinct probability-domain operators that yield identical student models under capacity constraints. We introduce an axiomatic definition of probability-domain softening operators based on ranking preservation, continuity, entropy monotonicity, identity, and boundary behavior, and show that multiple non-equivalent operator families satisfy these axioms. All learning-theoretic guarantees are shown to hold uniformly across this operator class, independent of implementation details. These results provide theoretical grounding for black-box teacher distillation, partial-access settings such as top-$k$ truncation and text-only outputs, and privacy-preserving model compression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。