arXiv:2601.13528cs.CRcs.AI2026-01被引 5

用安全模型生成数据,训练出能作恶的开源模型

Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs

  • 用不违规的提示词诱导安全模型输出有害内容
  • 使开源模型恢复约40%的危险能力缺口
  • 适合关注模型安全风险的研究者阅读

模型开发者在前沿模型中部署防护机制以防止滥用,例如通过分类器过滤危险输出。本文展示,即使经过严密防护的模型,仍可通过诱发攻击(elicitation attacks)在开源模型中激发有害能力。该攻击包含三个阶段:(i) 构建与目标有害任务相邻领域但不直接请求危险信息的提示;(ii) 从受保护的前沿模型获取这些提示的响应;(iii) 用这些提示-输出对微调开源模型。由于提示本身未被识别为危险,不会被安全机制拦截。我们在危险化学品合成与处理领域评估该攻击,结果表明,攻击可恢复基线开源模型与不受限前沿模型之间约40%的能力差距。进一步显示,攻击效果随前沿模型能力及生成微调数据量增加而提升。本工作揭示了仅靠输出层面防护难以应对生态级风险。

原文摘要 · Abstract (English)

Model developers implement safeguards in frontier models to prevent misuse, for example, by employing classifiers to filter dangerous outputs. In this work, we demonstrate that even robustly safeguarded models can be used to elicit harmful capabilities in open-source models through elicitation attacks. Our elicitation attacks consist of three stages: (i) constructing prompts in adjacent domains to a target harmful task that do not request dangerous information; (ii) obtaining responses to these prompts from safeguarded frontier models; (iii) fine-tuning open-source models on these prompt-output pairs. Since the requested prompts cannot be used to directly cause harm, they are not refused by frontier model safeguards. We evaluate these elicitation attacks within the domain of hazardous chemical synthesis and processing, and demonstrate that our attacks recover approximately 40% of the capability gap between the base open-source model and an unrestricted frontier model. We then show that the efficacy of elicitation attacks scales with the capability of the frontier model and the amount of generated fine-tuning data. Our work demonstrates the challenge of mitigating ecosystem level risks with output-level safeguards.

模型安全微调攻击有害能力防御漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。