提出零温度下安全不变性检测方法,确保生成内容无风险泄漏。
Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion
- 设计行为等价筛选器,对比目标模型与推测生成结果的安全表现。
- 48,072样本测试中,最大效应量仅0.024,远低于显著阈值。
- 适合关注推理加速中安全性的大模型应用开发者。
推测解码通过草稿模型预提案例供目标模型验证以加速推理,引发关键安全问题:在温度为零时,草稿侧行为是否会渗入经安全评分的输出?本文提出典型接受不变性筛选(TAIS),通过同一安全基准对齐目标模型与推测输出,要求字节级完全一致、双侧等效检验在±3个百分点内通过、每任务克里夫兰·赫系数| h |低于校准后0.1的虚无阈值。在包含16,783个验证样本及44,066个匹配扩展样本的测试中(fp16/bf16执行,标准与DPO对抗草稿、GPTQ-4bit草稿、两种子、四个安全基准),所有温度为零的vLLM堆栈在TAIS下均未检测到安全偏差。最大小组效应量仅为0.024,约为传统微小效应阈值的一阶;25/27项任务的TOST检验在±3pp边界通过(两项未通过为能力域沃尔德置信区间边缘情况,非真实不等价);DPO对抗草稿与标准草稿在4,006样本上字节完全一致;bf16改变36%-53%输出字节但未导致任一任务安全率越界。另一次700样本的70B生产规模探测(无匹配目标对照组,故不计为TAIS通过)显示,AdvBench拒绝率为0.839(95%威尔逊置信区间[0.809, 0.864])。本文不涉及采样温度、未测试框架、未测模型家族或如EAGLE、Medusa等树状推测变体。
原文摘要 · Abstract (English)
Speculative decoding accelerates inference by letting a draft model propose tokens for a target model to verify, raising a concrete safety question: at temperature zero, can draft-side behavior leak into safety-scored outputs? We answer with Typical-Acceptance Invariance Screen (TAIS), a behavioral-equivalence screen that pairs target-only and speculative outputs on the same safety battery and requires byte-identity evidence, TOST equivalence at +/-3pp, and per-task Cohen's h below a calibrated null cutoff of |h| < 0.1. Applied to a 16,783-sample confirmatory core plus 44,066 matched expansion samples (fp16/bf16 execution, canonical and DPO-adversarial drafts, GPTQ-4bit drafts, two seeds, and four safety benchmarks), the tested temperature-zero vLLM stacks show no detectable safety divergence under TAIS. The largest absolute Cohen's h on matched target-only versus speculative refusal is 0.024, roughly an order of magnitude below the conventional trivial-effect floor; 25 of 27 per-task TOST contrasts pass at the +/-3pp margin (the two non-pass contrasts are capability-domain Wald-CI edge cases at identical ceiling rates, not genuine non-equivalence); the DPO-adversarial draft produces byte-identical output to the canonical draft across 4,006 samples; and bf16 changes 36%-53% of output bytes without moving any per-task safety rate outside equivalence. A separate 4,006-sample 70B production-scale probe, which lacks a matched 70B target-only arm and is therefore not counted as a TAIS pass, produces AdvBench refusal 0.839 over 700 AdvBench completions with 95% Wilson CI [0.809, 0.864]. We make no claim about sampling temperatures, untested frameworks, untested model families, or tree-speculation variants such as EAGLE and Medusa.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。