arXiv:2608.17202cs.AIcs.CR2026-08

用伪造答案欺骗安全移除攻击,让模型在被篡改后仍输出看似合理实则错误的内容。

Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models

  • 在攻击模拟中训练伪回答,使模型被篡改后仍能生成流畅但虚假的回应。
  • 6个大模型测试中,被攻击后0.51-0.90的回答为伪造内容,防御提升0.27-0.84。
  • 适合关注开放权重模型安全防护的研究者与安全团队参考。

开放权重语言模型的安全对齐极易被移除:通过消融可在数分钟内从权重中提取拒绝响应方向,且目前所知的任何发布时防御均无法持久阻止。无法防止的攻击可被欺骗。本文提出防御机制“诱饵强化”(Fool's Gold),承认拒绝路径被剥离,并毒化其收益:一旦拒绝被移除,大多数对危险请求的回复均为自信、流畅但关键元素被篡改的伪答案。这些诱饵在攻击的可微模拟中训练,仅在被攻击状态下表达;拒绝锚点与良性约束保持原始状态行为。我们在五个模型族的七个模型(9B-122B,密集与专家混合)上实现该方法。六款通过预注册有效性门槛的模型中,攻击后对保留提示的回复中0.51-0.90为诱饵,其中0.27-0.84归因于防御;所有六款均在注册的良性行为与能力预算内;第七款(更小)未通过门槛(边界情况)。结果在冻结测试集或未触及分组中复现。主张为认识论性质:在无独立真实标签情况下,我们测试的所有观测表面均无法区分伪造答案与正确答案——在外部红队基准的CBRNE相关子集上,经防御的122B模型在匹配质量回答中犯错率达0.82-0.86,而未防御模型不超过0.10。重复采样也无法重建信任:在K=64时,元素级共识仅在0.083-0.625的提示中重构出可用程序,而未防御模型为0.58-0.96,且无无标签方式区分二者;最弱模型下该主张仅按单次抽样成立。评估涵盖化学与生物危害;该防御不处理上下文越狱,且仅保护初始发布的受保护权重。

原文摘要 · Abstract (English)

Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably. What cannot be prevented can be deceived. Our defense, decoy hardening ("Fool's Gold"), concedes the refusal strip and poisons its payoff: once refusal is stripped, most answers to hazardous operational requests are confident, fluent decoys whose critical elements are falsified. Decoys are trained inside a differentiable simulation of the attack, expressing only in the attacked state; a refusal pin and benign leash hold clean-state behavior to the original. We instantiate it on seven models from five families (9B-122B, dense and mixture-of-experts). On the six models passing our pre-registered efficacy gate, 0.51-0.90 of attacked-state responses to held-out prompts are decoys, +0.27-0.84 attributable to the defense; all six stay within registered benign-behavior and capability budgets; the seventh (smaller) fails the gate (boundary case). Rates replicate on a frozen test split or untouched strata. The claim is epistemic: without independent ground truth, no observation surface we tested separates falsified answers from correct ones - on external red-team benchmarks' CBRNE-adjacent slice, the defended 122B is fatally wrong on 0.82-0.86 of matched-quality answers vs at most 0.10 undefended. Repeated sampling does not restore trust: element-wise consensus at K=64 reconstructs a fully usable procedure on 0.083-0.625 of prompts where the instrument validates, vs 0.58-0.96 undefended, with no label-free way to tell the regimes apart; on the weakest such model the claim is per-draw only. We evaluate chemical and biological hazards; the defense does not address in-context jailbreaks and protects only the initially released defended weights.

模型安全对抗防御开放权重诱饵机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。