arXiv:2607.21804cs.CRcs.CL2026-07

提出新型对抗性后缀攻击,让推测解码失效却保持任务质量

Adversarial Prompts for Acceptance Collapse in Speculative Decoding

论文配图:Adversarial Prompts for Acceptance Collapse in Speculative Decoding
图 1 · 摘自论文原文
  • 设计软坍塌机制,诱导目标模型拒绝生成
  • 在GSM8K上使平均采样时间增加62.3%仍保任务正确率
  • 适用于多种模型与推理策略,揭示深层安全漏洞

无损加速方法如推测解码通过动态对齐草稿模型与目标模型实现显著推理提速。然而,这种语义等价性隐藏了严重操作漏洞:草稿-目标对齐可被系统性攻击。本文提出ADSD,据我们所知是首个通过将草稿概率质量推向目标模型难以接受的词元来导致验证器接受失败的提示后缀攻击。ADSD采用基于非对称推测接受规则的验证器对齐代理(Soft-Collapse),并引入目标保留目标函数以避免明显任务破坏。该攻击成功生成高度有效的对抗性后缀。在GSM8K数据集上,攻击使平均样本时间增加62.3%,同时保持任务质量。进一步表明此漏洞存在于不同领域、推测解码策略及模型架构中。

原文摘要 · Abstract (English)

Lossless acceleration schemes, such as speculative decoding, promise significant inference speedups by relying on dynamic token-level alignment between a draft and a target model. However, this guarantee of semantic equivalence masks a severe operational vulnerability: draft-target alignment can be systematically attacked. In this paper, we introduce ADSD, which, to the best of our knowledge, is the first prompt-suffix attack that collapses verifier acceptance by pushing draft probability mass toward tokens the target is unlikely to accept. ADSD uses Soft-Collapse, a verifier-aligned surrogate derived from the asymmetric speculative acceptance rule, together with a target-preservation objective that discourages obvious task corruption. ADSD successfully generates highly effective adversarial suffixes. On the GSM8K dataset, our attack increases the mean sample time by 62.3% while preserving the task quality. We further show that this vulnerability exists across different domains, speculative decoding strategies, and model architectures.

对抗攻击推理加速模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。