arXiv:2608.23375cs.CRcs.AI2026-08

对抗性提示破坏基于Gumbel的验证,让大模型漏出更多秘密信息。

Adversarial Entropy Inflation Against Gumbel-Based Inference Verification

  • 用恶意提示破坏语法结构,扩大可接受词元范围
  • 攻击使每令牌泄露比特翻倍,延迟降低至60x-118x
  • 防御阈值应动态根据本地熵调整,而非固定

基于Gumbel的推理验证通过仅允许源于诚实GPU非确定性的合理词元选择来限制大语言模型权重外泄,对良性提示流量下的隐蔽攻击者造成超过200倍的延迟。该防御假设攻击者为被动状态;我们发现若攻击者主动控制提示分布,防御效果将显著下降。由于验证器可接受词元集合大小由模型自身输出熵决定,针对语法和子词结构进行工程设计的恶意提示会扩大该集合,从而打开更大的隐蔽信道。在六个参数规模从1B到32B、三组随机种子的指令微调模型上,最强攻击(字符与语系级扰动)使每令牌泄露比特数接近翻倍,延迟因子降至60x–118x。结果表明,静态且基于良性流量校准的阈值不足以支撑此防御,抖动宽容阈值应动态依据局部词元熵进行校准。

原文摘要 · Abstract (English)

Gumbel-based inference verification bounds LLM weight exfiltration by only forgiving token choices that plausibly arise from honest GPU nondeterminism, reporting a >200x slowdown for a steganographic adversary under benign prompt traffic. This bound assumes a passive attacker; we show it degrades sharply against an adversary who instead controls the prompt distribution. Because the verifier's admissible-token-set size is driven by the model's own output entropy, prompts engineered to break grammatical and sub-word structure -- rather than benign conversational traffic -- widen that set and open a materially larger covert channel. Across six instruction-tuned models spanning 1B to 32B parameters and three random seeds, our strongest attack (character- and script-level disruption) roughly doubles bits leaked per token relative to benign prompts, cutting the slowdown factor to 60x - 118x. These results indicate that static, benign-traffic-calibrated thresholds are insufficient for this defense, and that jitter-forgiveness thresholds should instead be calibrated dynamically against local token entropy.

大模型安全对抗攻击隐写术熵分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。