让小模型复现大模型的学习环境,提升知识迁移效果
$\mathcal{X}$-KD: General Experiential Knowledge Distillation for Large Language Models
- 通过还原教师模型的原始学习环境进行知识迁移
- 在多个任务上超越基线方法,且更节省数据
- 适合追求高效、高质量模型压缩的研究者
大型语言模型的知识蒸馏日益重要。现有方法多关注模仿教师行为,却忽略了塑造教师知识的原始学习环境。受经验学习理论和逆强化学习启发,我们提出通用型经验知识蒸馏($Δ$-KD),使学生模型能在教师的原始学习环境中学习。$Δ$-KD采用近似变分奖励模仿学习(AVRIL)框架,联合建模教师原始奖励函数并执行策略蒸馏,促进学生策略与原奖励函数的一致性。推导表明,$Δ$-KD遵循监督学习框架,适用于序列级与基于差异的蒸馏方法,体现其简洁性与灵活性。实验显示,在摘要生成、机器翻译和算术推理任务中,$Δ$-KD优于广义蒸馏与MiniLLM基线。此外,其在性能与多样性权衡及数据效率方面也优于基线方法。
原文摘要 · Abstract (English)
Knowledge Distillation (KD) for Large Language Models (LLMs) has become increasingly important as models grow in size and complexity. While existing distillation approaches focus on imitating teacher behavior, they often overlook the original learning environment that shaped the teacher's knowledge. Inspired by the experiential learning theory and inverse reinforcement learning, we propose Experiential Knowledge Distillation ($\mathcal{X}$-KD), a novel and general framework that enables student models to learn in the teacher's original learning environment. $\mathcal{X}$-KD adopts the Approximated Variational Reward Imitation Learning (AVRIL) framework to jointly model the teacher's original reward function and perform policy distillation, encouraging consistency between the student policy and the original reward function. Our derivation demonstrates that $\mathcal{X}$-KD follows the supervised learning framework and applies to both sequence-level and divergence-based distillation methods, underlining the simplicity and flexibility of our approach. Empirical results show that $\mathcal{X}$-KD outperforms the generalized KD and MiniLLM baselines on abstractive summarization, machine translation, and arithmetic reasoning tasks. Additionally, $\mathcal{X}$-KD achieves better performance-diversity trade-off and data efficiency than baseline KD approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。