提出新框架PruneShift,评估结构化剪枝决策的可靠性。
PruneShift: A Framework for Evaluating Decision Reliability in Structured Pruning

- 分离预测精度、选择器附近拟合度与最终决策质量进行评估
- 实证发现7/20外部测试中剪枝选择更优,但多数结果不显著
- 适合关注剪枝方法可信度的研究者和实践者
结构化剪枝因直接评估所有可行掩码成本过高,通常依赖代理目标。现有评估多报告广泛采样掩码下的平均代理误差或等级相关性,但未直接检验代理所选掩码的性能。本文提出PruneShift框架,将评估分为三部分:全局预测保真度、选择器输出附近的保真度以及所选剪枝决策的质量。我们证明了斯皮尔曼与肯德尔一致性可趋近于1,而归一化选择遗憾仍可能达到最大值。进一步推导出基于均匀误差、选择器次优性、决策裕度、密度比和比较质量的充分条件,并给出有限池证书及显式超额成本上界。四个研究分别验证不同环节:外部TextbookQA显示20个同时区间中7个支持代理选择,6个支持固定对照组,7个跨零;固定Natural Questions池中4种设置仅1种严格提升;受控QQP实验在全部16个预设终点支持所提覆盖机制,但充分边界保守;在OPT-125M的受限OSSCAR重构中,局部保真度优于全局保真度的主端点达68/75。独立固定掩码确认在24/25端点中无结论,仅1个支持对照组。结果表明,预测拟合、决策可靠性和剪枝方法质量需独立证据。
原文摘要 · Abstract (English)
Structured pruning uses surrogate objectives because direct task evaluation over every feasible mask is too expensive. Most evaluations report average surrogate error or rank correlation on broadly sampled masks. These summaries do not directly test the mask chosen by the surrogate. We introduce PruneShift, an evaluation framework that separates broad predictive fidelity, fidelity near selector outputs, and the quality of the selected pruning decision. We first prove that Spearman and Kendall agreement can approach one while normalized selection regret remains maximal. We then derive sufficient conditions based on uniform error, selector suboptimality, decision margin, density ratio, and comparison mass. The analysis also yields a finite pool certificate with an explicit excess cost bound. Four studies test different links in this argument. External TextbookQA confirmation is heterogeneous: 7 of 20 simultaneous intervals favor the surrogate-selected mask, 6 favor its fixed comparator, and 7 cross zero. On a fixed Natural Questions pool, strict improvement holds in one of four settings. A controlled QQP experiment supports the proposed coverage mechanism in all 16 prespecified endpoints, although the sufficient bounds are conservative. Finally, a restricted OSSCAR reconstruction study on OPT-125M shows better local than broad fidelity in 68 of 75 primary endpoints. Independent fixed-mask confirmation is inconclusive in 24 of 25 endpoints and favors the comparator in one. These results show why predictive fit, decision reliability, and pruning method quality require separate evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。