arXiv:2602.15438cs.LGcs.AI2026-02

用对数几率距离提升模型表示相似性,让学生更好继承教师的可解释概念。

Logit Distance Bounds Representational Similarity

  • 基于对数几率差定义新距离度量,保证表示相似性
  • 实验证明该方法在图像与合成数据上保留更多可线性恢复的概念
  • 适合关注模型可解释性与知识迁移的研究者

对于包括自回归语言模型在内的广泛判别模型,可识别性结果表明:若两模型诱导相同的条件分布,则其内部表示在可逆线性变换下相等。我们探讨当分布接近而非完全相等时,是否仍能得出类似结论。基于Nielsen等人(2025)指出的KL散度接近不必然带来高线性表示相似性的观察,本文研究一种基于对数几率差异的分布距离,并证明该距离的接近性可导出线性相似性保证。具体而言,我们定义了基于模型可识别类的表示差异度量,并证明其被对数几率距离所控制。进一步表明,当模型概率远离零时,KL散度可上界对数几率距离;但该界在实践中无法提供非平凡控制。因此,基于KL的蒸馏虽能匹配教师预测,却可能丢失线性可恢复的人类可解释概念。在合成数据和图像数据集上的蒸馏实验显示,基于对数几率距离的蒸馏方法生成的学生模型具有更高的线性表示相似性,并更好地保留了教师的线性可恢复概念。

原文摘要 · Abstract (English)

For a broad family of discriminative models that includes autoregressive language models, identifiability results imply that if two models induce the same conditional distributions, then their internal representations are equal up to an invertible linear transformation. We ask whether an analogous conclusion holds approximately when the distributions are close instead of equal. Building on the observation of Nielsen et al. (2025) that closeness in KL divergence need not imply high linear representational similarity, we study a distributional distance based on logit differences and show that closeness in this distance does yield linear similarity guarantees. Specifically, we define a representational dissimilarity measure based on the models' identifiability class and prove that it is bounded by the logit distance. We further show that, when model probabilities are bounded away from zero, KL divergence upper-bounds logit distance; yet the resulting bound fails to provide nontrivial control in practice. As a consequence, KL-based distillation can match a teacher's predictions while failing to preserve linear representational properties, such as linear-probe recoverability of human-interpretable concepts. In distillation experiments on synthetic and image datasets, logit-distance distillation yields students with higher linear representational similarity and better preservation of the teacher's linearly recoverable concepts.

知识蒸馏表示学习可解释性对数几率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。