arXiv:2410.02220cs.CRcs.AI2024-10被引 3

通过智能数据筛选提升大模型抗越狱攻击能力

Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks

  • 动态筛选训练数据以增强模型防御力
  • 在微调全流程中实现100%安全响应生成
  • 适合关注模型安全的开发者与研究者

大语言模型通过微调实现定制化应用,但近期研究表明,恶意样本会破坏模型鲁棒性并放大有害行为,即越狱攻击。为应对该问题,我们提出一种自适应数据筛选方法,可将任意文本优化为有效对抗有害样本的资源。为避免引入额外防御模块,我们构建覆盖定制化全生命周期的综合缓解框架:微调前免疫未来越狱尝试,微调中实时化解风险,微调后恢复受损模型。实验表明,该方法显著降低越狱影响,实现高达100%的安全响应生成率。结合自适应数据筛选与全周期策略,本工作为降低越狱风险、保障大模型安全适配提供了坚实方案。

原文摘要 · Abstract (English)

Large language models (LLMs) are widely adapted for downstream applications through fine-tuning, a process named customization. However, recent studies have identified a vulnerability during this process, where malicious samples can compromise the robustness of LLMs and amplify harmful behaviors-an attack commonly referred to as jailbreaking. To address this challenge, we propose an adaptive data curation approach allowing any text to be curated to enhance its effectiveness in counteracting harmful samples during customization. To avoid the need for additional defensive modules, we further introduce a comprehensive mitigation framework spanning the lifecycle of the customization process: before customization to immunize LLMs against future jailbreak attempts, during customization to neutralize risks, and after customization to restore compromised models. Experimental results demonstrate a significant reduction in jailbreaking effects, achieving up to a 100% success rate in generating safe responses. By combining adaptive data curation with lifecycle-based mitigation strategies, this work represents a solid step forward in mitigating jailbreaking risks and ensuring the secure adaptation of LLMs.

大模型安全越狱攻击数据筛选微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。