为4比特量化测试设定可量化的可靠性门槛,避免结果被随机噪声误导。
Pre-Registering the Detectable Effect: A Paired-MDE Budget for 4-bit Quantization Benchmarks, with a Pilot Audit

- 用配对样本法推导出量化对比的最小可检测效应下限,指导实验设计
- 实测显示多数误差在统计噪音范围内,说明部分‘不可靠’源于抽样波动
- 建议先固定提示模板,否则其变异会混入量化评估的噪声中
本文提出一种面向4比特量化基准测试的预注册方法,将经典配对二项样本量计算(Miettinen, 1968)适配至量化场景,给出保守的最小可检测效应(MDE)边界:δ* ≤ (z₁₋α/₂ + z₁₋β)√(ρ_d/m),其中 m 为配对样本数,ρ_d 为FP16与NF4的不一致率。该边界将“量化结论是否可信”转化为可预注册的一行预算。我们在四个模型、四个基准(k=5个分组,每组n=100)上验证该边界,并补充了MMLU上的提示模板研究,比较量化噪声与提示噪声尺度。假设ρ_d=0.10(规划值),所有观测到的NF4-FP16差异均低于隐含的MDE;多数跨分组标准差在±1.5 pp内,接近二项分布参考值√(p(1−p)/n),表明大量所谓“基准不可靠性”实为抽样噪声。唯一临界单元(OPT-WinoGrande,|Δ|=3.2 pp)在ρ_d=0.10时低于MDE,但在ρ_d=0.05时高于MDE,凸显规划中的权衡。在MMLU上,提示模板范围2–10 pp覆盖最大量化差异(3.2 pp),表明未固定提示模板的审计会将模板变异纳入噪声底限。文中附五行预注册模板。
原文摘要 · Abstract (English)
This is a planning-method note with an unpaired pilot audit. We adapt the classical paired-binary sample-size calculation (Miettinen, 1968) to quantization benchmarks, giving a conservative minimum detectable effect (MDE) bound $δ^{*} \le (z_{1-α/2}+z_{1-β})\sqrt{ρ_d/m}$ in the paired item count $m$ and the FP16-NF4 disagreement rate $ρ_d$. The bound turns "how reliable is my quantization claim?" into a one-line budget a benchmark designer can commit to before running. We illustrate the bound on four models and four benchmarks ($k=5$ splits of $n=100$), and add a parallel MMLU prompt-template study to put the bound's quantization-noise scale alongside the prompt-noise scale. Assuming $ρ_d=0.10$ (an unmeasured planning value), all observed NF4-FP16 deltas fall below the implied MDE, and most cross-split SDs lie within $\pm 1.5$ pp of the binomial reference $\sqrt{p(1-p)/n}$, so much of the variance reported as "benchmark unreliability" on $n=100$ subsamples is binomial sampling noise. The single borderline cell (OPT-WinoGrande, $|Δ|=3.2$ pp) is below the implied MDE at $ρ_d=0.10$ but above it at $ρ_d=0.05$, illustrating the planning trade-off the bound makes explicit. On MMLU, prompt-template ranges of 2-10 pp meet or exceed the largest observed quantization delta (3.2 pp), so a quantization audit that does not first fix the prompt template absorbs template variance into its noise floor. We complement the bound with a five-line pre-registration template.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。