arXiv:2608.15215stat.MLcs.LG2026-08

用多温度分布视角改进知识蒸馏,让模型更懂错误类型。

The Distributional View of Knowledge Distillation

论文配图:The Distributional View of Knowledge Distillation
图 1 · 摘自论文原文
  • 将教师模型视为多个温度下的分布族,学生通过几何感知聚合学习
  • 实验发现提升效果取决于温度分散度而非视图数量,且运输聚合优于平均
  • 蒸馏效果好坏取决于教师与监督微调的差距大小,非固定损失函数

词元级知识蒸馏在每个位置匹配两个条件分布,但传统目标仅逐点比较:Kullback-Leibler梯度无法区分错误词元接收概率的位置。本文提出分布视角,将教师表示为一系列多温度视图(其对数温度路径的边缘分布),学生在基于嵌入的基代价下对抗这些视图的几何感知聚合进行训练。我们形式化了设计空间(混合、对数线性池化、熵正则Wasserstein均值,以及中心与路径形式的无偏Sinkhorn散度旗舰),证明对温度视图的对数线性池化等价于单一温度,并给出可检验预测的多边缘Schrodinger桥解读。在指令微调的Pythia对上,实验揭示三条经验规律:(i) 分散律——多温度聚合收益随视图有效温度分散度单调增长,而非视图数量;(ii) 分散视图解锁聚合算子——当运输聚合超越平均时,均值与算术混合开始分离;(iii) 双 regime 图景由天花板间隙 Γ = PPL_SFT - PPL_T 控制:当教师仅轻微优于监督微调时,温和运输目标最优,但蒸馏仍不及监督微调;而当达到真实天花板时,排名反转,保真度-泛化相关性符号翻转。我们主张‘最佳蒸馏损失’并非损失本身属性,而是 Γ 的函数。

原文摘要 · Abstract (English)

Token-level knowledge distillation (KD) matches two conditional distributions per position, yet the standard objectives compare them pointwise: a Kullback-Leibler gradient is blind to which wrong token receives probability mass. We develop a distributional view in which the teacher is represented not by a single softened output but by a family of multi-temperature views - marginals of the annealing path of its logits - and the student is trained against a geometry-aware aggregate of these views under an embedding-based ground cost. We formalize the resulting design space (mixtures, log-linear pooling, entropic Wasserstein barycenters, and a debiased Sinkhorn-divergence flagship in hub and path forms), prove an exact collapse result showing log-linear pooling of tempered views is equivalent to a single temperature, and give a multi-marginal Schrodinger-bridge reading that yields falsifiable predictions. On instruction-tuned Pythia pairs, experiments yield three empirical laws: (i) dispersion law - the benefit of multi-temperature aggregation grows monotonically with the effective temperature dispersion of the views, not with their number; (ii) dispersed views unlock the aggregation operator - the barycenter separates from the arithmetic mixture exactly when transport-based aggregation starts to beat averaging; and (iii) two-regime picture governed by the ceiling gap $Γ=\mathrm{PPL}_{\mathrm{SFT}}-\mathrm{PPL}_{T}$: when the fine-tuned teacher barely beats a supervised student the gentle transport objective is the best KD loss but no KD beats supervised fine-tuning, whereas at a real ceiling the ranking inverts - and the sign of the fidelity-generalization correlation flips. We argue that "which distillation loss is the best" is not a fixed property of the loss but a function of $Γ$.

知识蒸馏分布对齐生成模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。