arXiv:2505.10822cs.LG2025-05

揭秘知识蒸馏中模型内部如何重组与压缩,揭示学生模型依赖更少组件的机制。

Distilled Circuits: A Mechanistic Study of Internal Restructuring in Knowledge Distillation

  • 用可解释性方法分析教师与学生模型的内部计算差异。
  • 学生模型会重组、压缩甚至丢弃教师组件,依赖更少核心模块。
  • 提出基于影响权重的对齐度量,超越输出相似性评估功能一致性。

知识蒸馏通过训练小型学生模型模仿大型教师模型的输出,实现模型压缩。然而,这一过程中的内部计算转变机制仍不清晰。本文采用机制可解释性技术,分析教师模型(如GPT2)与学生模型(如DistilGPT2)在内部电路、表示和激活模式上的差异,并将结论推广至双向架构及更大模型对。研究发现,学生模型会重新组织、压缩并舍弃教师组件,通常表现出对少数核心组件更强的依赖。为量化输出之外的功能对齐,本文提出一种基于影响权重的组件相似性度量,经多任务验证有效。结果表明,尽管蒸馏保留了整体功能行为,但内部计算方式发生显著变化,对蒸馏模型的鲁棒性和泛化能力具有重要影响。

原文摘要 · Abstract (English)

Knowledge distillation compresses a larger neural model (teacher) into smaller, faster student models by training them to match teacher outputs. However, the internal computational transformations that occur during this process remain poorly understood. We apply techniques from mechanistic interpretability to analyze how internal circuits, representations, and activation patterns differ between teachers and students. Focusing on GPT2 and its distilled counterpart DistilGPT2, and generalizing our findings to both bidirectional architectures and larger model pairs, we find that student models can reorganize, compress, and discard teacher components, often resulting in a stronger reliance on fewer individual components. To quantify functional alignment beyond output similarity, we introduce an alignment metric based on influence-weighted component similarity, validated across multiple tasks. Our findings reveal that while knowledge distillation preserves broad functional behaviors, it also causes significant shifts in internal computation, with important implications for the robustness and generalization capacity of distilled models.

知识蒸馏模型压缩可解释性机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。