让学生模型学会老师的真实计算逻辑,而非仅模仿输出。
Circuit Distillation
- 通过匹配师生模型中功能对应的电路组件,对齐内部表示。
- 在实体追踪和心智理论任务中,仅调整少量参数就超越传统蒸馏。
- 适合需要可解释、可控能力迁移的研究者使用。
模型蒸馏通常关注行为模仿,即让学生模型复制教师模型的输出,而将教师的内部计算视为黑箱。本文提出一种新方法:蒸馏教师模型所实现的底层计算机制。具体地,提出电路蒸馏,引入目标函数以对齐教师与学生模型中对应电路组件的内部表征。我们提出一种方法来匹配“功能对应”的电路组件,并引入损失函数以反映这些组件所诱导表征的相似性。我们在基于 Llama3 系列模型的实体追踪和心智理论(ToM)任务上评估了该方法。结果表明,电路蒸馏优于标准蒸馏,在仅调整学生模型少量参数的情况下成功转移了算法能力。本工作证明了机制转移的可行性,或可实现通过可解释且可控的学生内部机制,高效蒸馏教师的特定能力。
原文摘要 · Abstract (English)
Model distillation typically focuses on behavioral mimicry, where a student model is trained to replicate a teacher's output while treating its internal computations as a black box. In this work we propose an alternative approach: Distilling the underlying computational mechanisms implemented by a teacher model. Specifically, we propose circuit distillation, which introduces an objective to align internal representations between analogous circuit components in teacher and student models. We propose a method to match ``functionally correspondent'' circuit components and introduce a loss reflecting similarities between the representations that these induce. We evaluate circuit distillation on entity tracking and theory of mind (ToM) tasks using models from the Llama3 family. Our results demonstrate that circuit distillation outperforms standard distillation, successfully transferring algorithmic capabilities by adjusting only a small, targeted subset of student model parameters. This work establishes the feasibility of transferring mechanisms, which may in turn allow for efficient distillation of targeted teacher capabilities via interpretable and controllable internal student mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。