arXiv:2510.10618cs.CL2025-10NeurIPS被引 9

优化校准数据可更好保留大模型压缩后的推理能力

Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization

  • 从激活模式分析校准数据影响,发现代表性和多样性是关键
  • 新方法在数学和代码生成任务上显著提升压缩后性能
  • 适合关注模型压缩中能力保持的研究者与工程师

后训练压缩是缩小大语言模型(LLM)以实现高效推理的常用方法。在剪枝、量化等压缩技术中,校准数据通过提供权重重要性与激活动态范围信息起关键作用。然而,校准数据对压缩后模型能力的影响尚不明确。现有研究多局限于数据来源或样本量对语言建模或常识推理性能的影响,缺乏系统性分析,尤其在组合性质和领域匹配度方面。本文旨在填补这一空白,从激活模式角度深入分析其内在机制,重点考察数学求解与代码生成等高级复杂推理能力。研究发现,校准数据在激活空间中的代表性和多样性更根本地决定其质量。基于此,我们提出一种校准数据筛选框架,显著提升了现有后训练压缩方法在保留关键大模型能力方面的表现。代码已开源。

原文摘要 · Abstract (English)

Post-training compression has been a widely employed approach to scale down large language model (LLM) and facilitate efficient inference. In various proposed compression methods, including pruning and quantization, calibration data plays a vital role by informing the weight importance and activation dynamic ranges. However, how calibration data impacts the LLM capability after compression is less explored. Few of the existing works, though recognizing the significance of this study, only investigate the language modeling or commonsense reasoning performance degradation from limited angles, like the data sources or sample amounts. More systematic research is still needed to examine the impacts on different LLM capabilities in terms of compositional properties and domain correspondence of calibration data. In this work, we aim at bridging this gap and further analyze underlying influencing mechanisms from the activation pattern perspective. Especially, we explore the calibration data's impacts on high-level complex reasoning capabilities, like math problem solving and code generation. Delving into the underlying mechanism, we find that the representativeness and diversity in activation space more fundamentally determine the quality of calibration data. Finally, we propose a calibration data curation framework based on such observations and analysis, enhancing the performance of existing post-training compression methods on preserving critical LLM capabilities. Our code is provided in \href{https://github.com/BokwaiHo/COLA.git}{Link}.

大模型压缩校准数据推理能力保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。