arXiv:2506.10020cs.CRcs.AI2025-06

用攻击技巧生成安全训练数据,让大模型更抗恶意指令

From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment

  • 通过检测模型拒绝信号,动态注入特定短语诱导输出有害内容
  • 在4个基准上将有害响应率从2.15%提升至61.04%的峰值
  • 生成的数据能增强模型抗风险能力,适合安全对齐研究者使用

大语言模型的安全对齐通常依赖大量人工标注的偏好数据,成本高昂。尽管合成数据是潜在替代方案,但现有方法多依赖复杂的迭代提示或辅助模型。为此,我们提出无需训练、通用且简单的拒绝感知自适应注入(RAAI)框架,将大模型攻击技术转化为安全对齐工具。RAAI通过识别内部拒绝信号,动态注入预设短语以诱导出有害但流畅的回复。实验表明,RAAI可有效越狱大模型,在四个基准上平均将有害响应率从基线2.15%提升至61.04%。更重要的是,使用RAAI生成的合成数据微调模型后,其在对抗有害提示时表现更稳健,同时保持在MMLU和ARC等标准任务上的通用能力。本工作揭示了大模型攻击方法可被重构为可扩展、可控的安全对齐实用工具。

原文摘要 · Abstract (English)

Safely aligning large language models (LLMs) often demands extensive human-labeled preference data, a process that's both costly and time-consuming. While synthetic data offers a promising alternative, current methods frequently rely on complex iterative prompting or auxiliary models. To address this, we introduce Refusal-Aware Adaptive Injection (RAAI), a straightforward, training-free, and model-agnostic framework that repurposes LLM attack techniques. RAAI works by detecting internal refusal signals and adaptively injecting predefined phrases to elicit harmful, yet fluent, completions. Our experiments show RAAI effectively jailbreaks LLMs, increasing the harmful response rate from a baseline of 2.15% to up to 61.04% on average across four benchmarks. Crucially, fine-tuning LLMs with the synthetic data generated by RAAI improves model robustness against harmful prompts while preserving general capabilities on standard tasks like MMLU and ARC. This work highlights how LLM attack methodologies can be reframed as practical tools for scalable and controllable safety alignment.

安全对齐模型攻击合成数据大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。