arXiv:2604.15794cs.LGcs.AI2026-04

通过自蒸馏修复大模型压缩和遗忘导致的性能下降

Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting

  • 用自蒸馏微调恢复压缩/微调后的模型能力
  • 性能回升与隐藏层激活轨迹对齐度强相关
  • 适合研究模型退化机制或优化方法的学者

大语言模型在各类AI应用中表现卓越,但常因监督微调中的灾难性遗忘、量化和剪枝导致性能下降。本文提出基于自蒸馏微调(SDFT)的性能恢复框架,有效恢复模型能力。理论层面,我们指出模型生成能力依赖于隐藏层构建的高维流形结构。通过中心核对齐(CKA)量化学生与教师激活轨迹的对齐程度,利用其对正交变换和缩放的不变性,实验表明性能恢复与流形对齐度存在强相关性,证实自蒸馏能将学生模型的高维流形对齐至教师最优结构。本研究连接了实践恢复框架与几何表示理论,揭示了自蒸馏的内在机理。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable success, underpinning diverse AI applications. However, they often suffer from performance degradation due to factors such as catastrophic forgetting during Supervised Fine-Tuning (SFT), quantization, and pruning. In this work, we introduce a performance recovery framework based on Self-Distillation Fine-Tuning (SDFT) that effectively restores model capabilities. Complementing this practical contribution, we provide a rigorous theoretical explanation for the underlying recovery mechanism. We posit that an LLM's generative capability fundamentally relies on the high-dimensional manifold constructed by its hidden layers. To investigate this, we employ Centered Kernel Alignment (CKA) to quantify the alignment between student and teacher activation trajectories, leveraging its invariance to orthogonal transformations and scaling. Our experiments demonstrate a strong correlation between performance recovery and manifold alignment, substantiating the claim that self-distillation effectively aligns the student's high-dimensional manifold with the optimal structure represented by the teacher. This study bridges the gap between practical recovery frameworks and geometric representation theory, offering new insights into the internal mechanisms of self-distillation.

自蒸馏模型压缩性能恢复流形学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。