压缩大模型同时保留推理与语言能力,双策略协同成新趋势
Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
- 结合知识蒸馏与数据蒸馏,用优化匹配和生成合成压缩模型
- 可使小模型保持大模型的推理与语言多样性,支持高效部署
- 适合关注大模型轻量化、跨领域应用的研究者与工程师
大型语言模型(LLMs)的指数级增长持续凸显其在计算与数据需求上的挑战。本文综述了两种互补范式:知识蒸馏(KD)与数据蒸馏(DD),均旨在压缩大模型的同时保留其高级推理能力与语言多样性。文章分析了KD的关键方法,如任务特定对齐、基于推理的训练与多教师框架,以及通过基于梯度匹配、隐空间正则化与生成合成实现的DD技术。在此基础上,探讨了融合KD与DD以构建更高效、可扩展的压缩策略。这些方法应对了模型可扩展性、架构异构性及涌现能力保持等长期挑战,并在医疗、教育等领域实现高效部署。尽管进展显著,仍面临保留涌现推理与语言多样性、适应动态教师模型与数据集、建立全面评估体系等挑战。本文通过整合方法创新、理论基础与实践洞见,为实现资源高效的大模型发展路径指明方向。
原文摘要 · Abstract (English)
The exponential growth of Large Language Models (LLMs) continues to highlight the need for efficient strategies to meet ever-expanding computational and data demands. This survey provides a comprehensive analysis of two complementary paradigms: Knowledge Distillation (KD) and Dataset Distillation (DD), both aimed at compressing LLMs while preserving their advanced reasoning capabilities and linguistic diversity. We first examine key methodologies in KD, such as task-specific alignment, rationale-based training, and multi-teacher frameworks, alongside DD techniques that synthesize compact, high-impact datasets through optimization-based gradient matching, latent space regularization, and generative synthesis. Building on these foundations, we explore how integrating KD and DD can produce more effective and scalable compression strategies. Together, these approaches address persistent challenges in model scalability, architectural heterogeneity, and the preservation of emergent LLM abilities. We further highlight applications across domains such as healthcare and education, where distillation enables efficient deployment without sacrificing performance. Despite substantial progress, open challenges remain in preserving emergent reasoning and linguistic diversity, enabling efficient adaptation to continually evolving teacher models and datasets, and establishing comprehensive evaluation protocols. By synthesizing methodological innovations, theoretical foundations, and practical insights, our survey charts a path toward sustainable, resource-efficient LLMs through the tighter integration of KD and DD principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。