arXiv:2603.22355stat.MLcs.CL2026-03被引 3

解析大模型低秩知识蒸馏的理论机制,给出收敛、泛化和信息保留的数学保障。

Demystifying Low-Rank Knowledge Distillation in Large Language Models: Convergence, Generalization, and Information-Theoretic Guarantees

  • 基于低秩投影建立理论框架,证明优化过程可保持不变,收敛速度为 $O(1/\\/sqrt{T})$。
  • 推导泛化误差界,揭示压缩率与泛化能力的权衡关系,误差随秩 $r$ 增长为 $O(r(m+n)/\\/sqrt{n})$。
  • 提出最优秩选择建议 $r^* = O(\\/sqrt{n})$,实验验证其在标准任务上有效匹配理论预测。

知识蒸馏已成为将大型语言模型压缩为高效部署架构的强大技术,同时保留其先进能力。近期低秩知识蒸馏方法(如低秩克隆)在实践中表现出色,仅用少量训练数据和计算开销即可达到全参数蒸馏的性能。然而这些方法的理论基础仍不清晰。本文建立了语言模型中低秩知识蒸馏的严格理论框架。在温和假设下,证明低秩投影能保持优化动态,给出显式的收敛速率 $O(1/\sqrt{T})$。推导出泛化边界,刻画模型压缩与泛化能力之间的根本权衡,表明泛化误差随秩参数以 $O(r(m+n)/\sqrt{n})$ 速率增长。进一步提供激活克隆机制的信息论分析,揭示其通过最大化教师与学生中间表示间的互信息来提升表达一致性。理论结果为秩的选择提供了原则性指导,数学上建议最优秩 $r^* = O(\sqrt{n})$,其中 $n$ 为样本量。在标准语言建模基准上的实验验证了理论预测,表明经验收敛性、秩缩放及泛化行为均与理论边界高度一致。

原文摘要 · Abstract (English)

Knowledge distillation has emerged as a powerful technique for compressing large language models (LLMs) into efficient, deployable architectures while preserving their advanced capabilities. Recent advances in low-rank knowledge distillation, particularly methods like Low-Rank Clone (LRC), have demonstrated remarkable empirical success, achieving comparable performance to full-parameter distillation with significantly reduced training data and computational overhead. However, the theoretical foundations underlying these methods remain poorly understood. In this paper, we establish a rigorous theoretical framework for low-rank knowledge distillation in language models. We prove that under mild assumptions, low-rank projection preserves the optimization dynamics, yielding explicit convergence rates of $O(1/\sqrt{T})$. We derive generalization bounds that characterize the fundamental trade-off between model compression and generalization capability, showing that the generalization error scales with the rank parameter as $O(r(m+n)/\sqrt{n})$. Furthermore, we provide an information-theoretic analysis of the activation cloning mechanism, revealing its role in maximizing the mutual information between the teacher's and student's intermediate representations. Our theoretical results offer principled guidelines for rank selection, mathematically suggesting an optimal rank $r^* = O(\sqrt{n})$ where $n$ is the sample size. Experimental validation on standard language modeling benchmarks confirms our theoretical predictions, demonstrating that the empirical convergence, rank scaling, and generalization behaviors align closely with our bounds.

知识蒸馏大模型压缩理论分析低秩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。