剪枝模型宽度发现:越剪知识越差,但指令跟随反而变强。
Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2
- 按峰值幅度剪枝GLU-MLP层,调节扩张比
- 指令遵循提升4.8分(+46%),数学推理保持稳定
- 适合关注模型压缩与能力平衡的研究者
对Llama-3.2模型中的GLU-MLP层进行结构化宽度剪枝,基于峰值幅度(PPM)准则,揭示了降低扩张比对不同模型能力的影响存在系统性二分现象。随着扩张比下降,依赖参数化知识的任务(如MMLU、GSM8K)性能和困惑度指标呈可预测下降;而指令跟随能力在2.4倍均衡比下显著提升(IFEval:Llama-3.2-1B +4.8分/+46%,Llama-3.2-3B +3.7分/+39%),多步推理能力也保持稳健(MUSR)。该现象在两种模型规模上均一致,挑战了压缩研究中剪枝导致均匀退化的主流假设。通过七个扩张比配置的综合基准评估,验证了扩张比是关键架构参数,能选择性重塑模型任务表现,而非仅作压缩指标。
原文摘要 · Abstract (English)
Structured width pruning of GLU-MLP layers in Llama-3.2 models, guided by the Peak-to-Peak Magnitude (PPM) criterion, reveals a systematic dichotomy in how reducing the expansion ratio affects different model capabilities. While performance on tasks relying on parametric knowledge (e.g., MMLU, GSM8K) and perplexity metrics degrades predictably with decreasing expansion ratios, instruction-following capabilities improve at the 2.4x equilibrium ratio (IFEval: +4.8 points / +46% in Llama-3.2-1B and +3.7 points / +39% in Llama-3.2-3B), and multi-step reasoning remains robust (MUSR). This pattern, observed consistently across both evaluated model sizes, challenges the prevailing assumption in compression research that pruning induces uniform degradation. To investigate this, we evaluated seven expansion ratio configurations using comprehensive benchmark suites that assess factual knowledge, mathematical reasoning, language comprehension, instruction-following, and truthfulness. Our analysis identifies the expansion ratio as a critical architectural parameter that selectively reshapes the model's task performance profile, rather than merely serving as a compression metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。