压缩模型等效性验证需关注个体差异,仅看平均精度易误导。
Certifying Compressed Language Models: An Audit and a Statistical Toolkit
- 用配对检验和声明阈值替代单一精度差值判断等效性
- 17个声称等效的模型中16个缺乏可验证的逐项输出数据
- 提出五行报告标准,强调公开逐项结果与实证依据
当前压缩模型等效性的判断常依赖微小的基准准确率差异,但此指标在模型最相似时最不具信息量——净误差是正负项抵消后的残余,而抵消在等效性宣称区间最彻底。分析1707组公开的模型-任务配对(13亿至4050亿参数)发现,净准确率差仅为整体变化率的五分之一;即使评分相同,个体样本仍存在显著分歧。对三个来源(方法论文、模型卡、厂商文档)中17个等效性主张进行预注册审计:16个符合资格,但无一声明数值等效边界,也未提供任务匹配的逐项输出,仅3个发布其他任务输出;5个数据不足无法量化评估。本文不评判真假,而是审计证据充分性,并提供缺失工具:在声明阈值下执行配对等效检验,认证表基于压缩后观察到的分歧计算,而非独立二项分布假设。受控实验显示,使用相同校准样本,GPTQ与AWQ在五次种子测试中,有5/8的确认单元因校准抽样变化导致方法排序反转。建议报告标准为五行:声明阈值、执行配对检验、报告变化率与净差、注明满足的样本量、公开逐项输出。该标准适用于任何值得比较的相近模型。所有逐项输出、协议与代码均已公开。
原文摘要 · Abstract (English)
A fraction of a point of benchmark accuracy is the usual evidence that a compressed model is equivalent to its original. That quantity is least informative when two models are most alike: a net delta is what survives cancellation between opposing per-item changes, and cancellation is most complete in the regime equivalence claims occupy. Across an atlas of 1,707 paired model-by-task cells mined from public per-item evaluation dumps (1.3B-405B), churn runs roughly five times the net accuracy delta, and cells scoring identically to their baseline still disagree on individual items. In a preregistered audit of 17 equivalence claims from three registered frames (method papers, model cards, vendor documentation), 16 are eligible. None states a prospective numerical equivalence margin, and none releases task-matched per-item outputs, though 3 release outputs for other tasks only; 5 report too little to assess numerically, so a reader cannot check them at any sample size. We audit evidential sufficiency, not truth: no claim is called false. We supply the missing instrument: paired equivalence testing at a declared margin, with certification tables giving the items an evaluation needs, computed from disagreement observed under compression, not from independent-binomial variance. A controlled experiment pairs GPTQ and AWQ on byte-identical calibration samples across five seeds. Under the frozen eight-cell decision rule H3 is supported: changing the calibration draw was sufficient to reverse the observed method ordering in 5 of 8 confirmatory cells. The reporting standard we propose is five lines: declare a margin, run the paired test, report churn beside net delta, cite the sample size you met, release per-item outputs. It applies to any comparison between two models alike enough to be worth comparing. All per-item outputs, protocols and code are released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。