arXiv:2606.24747cs.AIcs.CE2026-06

提出领域专用大模型压缩的量化规律,指导高效降本。

Scaling Laws for Task-Specific LLM Distillation

论文配图:Scaling Laws for Task-Specific LLM Distillation
图 1 · 摘自论文原文
  • 基于金融领域数据,研究压缩比与数据量对模型性能的影响
  • 链式思维监督可恢复剪枝丢失的通用知识,提升压缩后表现
  • 适用于需要低延迟、低成本部署的垂直领域模型优化

大语言模型在多领域表现优异,但其规模带来部署成本与延迟挑战。本文针对特定领域(以量化金融为例)的模型压缩,推导出经验性缩放定律,量化分析了数据集大小、压缩比、监督格式及迭代剪枝策略对领域内性能与通用知识表现的影响。通过对比基于输出概率和LoRA的蒸馏方法,引入混合链式思维监督损失,稳定了推理轨迹的KL散度蒸馏过程。结果表明:领域任务性能随压缩可预测下降,而通用知识基准在相同压缩点前大幅崩溃;监督格式是该权衡的关键因素,链式思维监督能有效恢复被剪枝抹去的通用知识。论文发布金融新闻数据集FinHeadlineMix、缩放定律结果与实用建议,构建可复用的领域专用压缩决策框架。

原文摘要 · Abstract (English)

Large Language Models (LLMs) achieve strong performance across a growing range of domains, yet their scale poses deployment challenges in applications where latency and cost constraints are critical. This paper derives empirical scaling laws for domain-specific LLM compression, quantifying how in-domain and general knowledge performance scale with dataset size, compression ratio, supervision format, and iterative pruning schedule. Using quantitative finance as our application domain, we compare logit-based and LoRA-based distillation under iterative structural pruning, introducing a blended chain-of-thought supervision loss that stabilizes KL-divergence distillation over reasoning traces. In-domain task quality degrades predictably under compression while general-knowledge benchmarks collapse well before the same point; supervision format is the key driver of this tradeoff, with chain-of-thought supervision actively recovering general knowledge that pruning erases. We release the headline dataset FinHeadlineMix, scaling law results, and practical recommendations to provide a reusable framework for domain-specific compression decisions.

大模型压缩缩放定律金融AI知识保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。