arXiv:2603.16440cs.LGcs.CL2026-03

按模型组件的功能重要性分配压缩预算,提升大模型压缩的可解释性。

Capability-Guided Compression: Toward Interpretability-Aware Budget Allocation for Large Language Models

  • 用稀疏自编码器生成能力密度图,指导不同模块差异化压缩
  • 高能力密度组件在更低压缩比下就出现性能突变,可提前预测
  • 能力密度与现有重要性指标无关,是全新压缩信号

大语言模型压缩虽通过剪枝、量化和低秩分解取得进展,但所有方法均存在根本缺陷:压缩预算分配未考虑各组件的功能编码信息。我们称之为‘能力盲压缩’问题,并指出这是导致困惑度评估对推理能力下降不敏感、以及近期马等(2026)描述的性能突变现象的根本原因。本文提出能力引导压缩(CGC),利用稀疏自编码器(SAE)生成的能力密度图,对Transformer组件进行差异化压缩预算分配。能力密度是结合特征广度、激活熵与跨输入一致性的形式化标量度量。理论上证明,能力密度高的组件结构冗余更低,且在更低压缩比下达到个体性能突变点,首次实现组件级突变预测。GPT-2 Medium实验表明,能力密度与Wanda重要性评分统计无关(斯皮尔曼相关系数 = -0.054,n=384),证实其为独立于现有度量的新压缩信号。尽管在困惑度基准上未见正向结果,但提供了对GPT-2 Medium作为测试基线不足的严谨诊断。理论框架、密度定义与正交性发现为能力感知压缩研究奠定基础。

原文摘要 · Abstract (English)

Large language model compression has made substantial progress through pruning, quantization, and low-rank decomposition, yet a fundamental limitation persists across all existing methods: compression budgets are allocated without any representation of what individual model components functionally encode. We term this the capability-blind compression problem and argue it is a root cause of two well-documented failures -- the insensitivity of perplexity-based evaluation to reasoning capability loss, and the abrupt phase transitions in model performance recently characterized by Ma et al. (2026). We propose Capability-Guided Compression (CGC), a framework that addresses this by using Sparse Autoencoder (SAE)-derived capability density maps to allocate differential compression budgets across transformer components. Capability density is a formally defined scalar measure combining the feature breadth, activation entropy, and cross-input consistency of a component's SAE feature activation distribution. We prove theoretically that components with higher capability density exhibit lower structural redundancy and reach their individual phase transition points at lower compression ratios, providing the first pre-compression mechanism for component-level phase transition prediction. Experiments on GPT-2 Medium confirm that capability density is statistically independent of Wanda importance scores (Spearman rho = -0.054, n = 384 heads), establishing it as a genuinely novel compression signal orthogonal to all existing importance metrics. We report a negative result on PPL-based compression comparison and provide a principled diagnosis identifying GPT-2 Medium as an insufficient test bed for the full CGC hypothesis. The theoretical framework, density formalism, and orthogonality finding constitute a foundation for capability-aware compression research.

模型压缩可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。