arXiv:2608.02940cs.AI2026-08

传统压缩评分可能选错模型,新方法确保最差群体表现最优。

When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning

  • 用信息界面视角分析压缩统计量的可靠性边界
  • 在三个大模型上降低最差组困惑度12.6%~20.9%
  • 适合关注群体公平性的模型压缩研究者

稳定的压缩评分仍可能选出更差的模型。在密集实验中,一个分半可靠路径二次评分预测16.1%收益,但选定端点却比两个对照组差6.0%–7.7%。当部署关注最差群体时,压缩统计量能证明什么?我们将每个统计量视为信息接口,其观测留下兼容端点风险表的纤维,仅在该纤维内一致的排序才可识别。锥与纤维身份量化剩余不确定性,匹配观测可反转汇总矩、群体局部矩和参考路径曲率的端点顺序。序列组合引入一个状态变量:各群体风险到当前最大值的松弛量。该向量决定所有无约束一步响应,边际条件保持路径中活动群体固定,相对漂移有界。实验遵循相同阶梯:在三个密集LLM上,早期保留分配使最差组困惑度增长减少12.6%–20.9%;目标匹配完整菜单选择优于参考方案2.7%–8.0%。在OLMoE全部16个路由层中,汇总端点刷新使保留最差组教师KL降低15.8%,优于最佳静态评分。计算匹配硬最大轨迹比汇总结果差32.7%,且两种自适应轨迹均未改善过量NLL。局部证据可缩小候选菜单,完整端点对菜单排序,多步主张还需控制演化中的活跃面与未来候选。

原文摘要 · Abstract (English)

A stable compression score can still select the worse model. In our dense study, a split-half reliable path-quadratic score predicted a 16.1\% gain, while the selected endpoints were 6.0--7.7% worse than two controls. We ask what a compression statistic can justify when deployment cares about the worst supplied group. We treat each statistic as an information interface. Its observation leaves a fiber of compatible endpoint-risk tables, and only orders fixed across that fiber are identified. Cone and fiber identities quantify the remaining uncertainty, while matched observations reverse endpoint order for pooled moments, group-local moments, and reference-path curvature. Sequential composition adds one state variable: the slack from each group risk to the current maximum. This vector determines every unrestricted one-step response, and a margin condition keeps the active group fixed along paths with bounded relative drift. The experiments follow the same ladder. Across three dense LLMs, an early-preserving allocation reduces worst-group perplexity inflation by 12.6--20.9%; target-matched complete-menu selection improves over its references by 2.7--8.0%. Across all 16 routed layers of OLMoE, pooled endpoint refresh lowers held-out worst-group teacher KL by 15.8% over the best static score. A compute-matched hard-max trajectory ends 32.7% worse than pooled, and neither adaptive trajectory improves excess NLL. Local evidence can narrow a menu. Complete endpoints rank that menu, while multistep claims also require control of the evolving active face and future candidates.

模型剪枝群体鲁棒性压缩评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。