剪枝不破坏预测可靠性,还能缩小置信集。
Calibration-Preserving Pruning: Compression as a Reliability Contract
- 用梯度显著性增强剪枝评分,保持校准性
- 50%稀疏下DBpedia任务置信集平均缩小至8.6
- 适合对预测可靠性要求高的分类场景
在独立于分层校准数据集的前提下,分裂共形预测可保证有限样本的边际覆盖。本文研究剪枝效率问题:能否在剪枝后仍保持得分几何结构以获得更小的有效预测集?提出校准保持剪枝(CPP),通过非共形梯度显著性扩展基础剪枝评分,并使用互斥的剪枝、验证选择、共形校准与测试分割。有界得分扰动意味着有界共形分位数偏移和可控集合膨胀,但该性质并非仅限于CPP。五组种子实验中,Qwen2.5-1.5B在50%稀疏下的大标签任务表现最佳。在DBpedia-14上,CPP-SparseGPT将平均置信集大小从10.1降至8.6,准确率由0.347升至0.366;CPP-Wanda将11.2降至9.0,准确率从0.310微降至0.295。在15个数据集-稀疏度组合中,CPP-SparseGPT在13个中产生更小集合,在11个中获得更高准确率。对照实验表明,通用监督梯度解释了大部分增益:真实标签的CPP与匹配的Wanda+SNIP无统计差异;而阈值感知候选标签的CPP达到7.8的平均集合大小,同时具备显式准确率和离线计算成本优势。RoBERTa-base与Llama-3-8B诊断支持迁移能力,但本文结论仅限于可靠性敏感的分类任务。
原文摘要 · Abstract (English)
Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal calibration split. We study the separate efficiency problem: can pruning preserve score geometry well enough to obtain smaller valid prediction sets? Calibration-Preserving Pruning (CPP) augments a base pruning score with nonconformity-gradient saliency and uses disjoint pruning, validation-selection, conformal-calibration, and test splits. Bounded score perturbations imply bounded conformal-quantile shifts and controlled set inflation, but do not make the generic coverage theorem CPP-specific. Final five-seed Qwen2.5-1.5B results at 50\% sparsity show the largest gains on large-label tasks. On DBpedia-14, CPP-SparseGPT reduces mean set size from \(10.1\) to \(8.6\) while changing accuracy from \(0.347\) to \(0.366\); CPP-Wanda reduces \(11.2\) to \(9.0\) with an accuracy trade-off from \(0.310\) to \(0.295\). Across 15 dataset--sparsity cells, CPP-SparseGPT produces smaller sets in 13 and higher accuracy in 11. Matched controls show that generic supervised gradients explain much of the gain: true-label CPP is not statistically resolved from matched Wanda+SNIP, whereas threshold-aware candidate-label CPP reaches \(7.8\) mean set size at explicit accuracy and offline-compute costs. RoBERTa-base and Llama-3-8B diagnostics support transfer, but our claims remain limited to reliability-sensitive classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。