通过多维攻防数据提升大模型生成安全性
A Method for Enhancing the Safety of Large Model Generation Based on Multi-dimensional Attack and Defense
- 构建多维度攻击指令与安全响应数据,增强模型对复杂攻击的防御能力
- 在Llama3.2上实验,显著提升模型在复杂指令攻击下的安全生成表现
- 兼顾安全性与通用能力,适合需要高可靠性的大模型应用
当前大模型在面对复杂攻击指令时容易生成有害内容,防御能力不足。本文提出一种基于多维攻防对齐数据的方法,以增强大模型的生成安全性。核心在于通过创新性增加攻击指令维度多样性与安全响应生成准确性,提升模型安全对齐学习效果。为验证方法有效性,除使用现有安全评估基准外,还设计了新的安全评估基准,并以Llama3.2为基线模型进行对比实验。结果表明,该方法在复杂指令攻击下显著提升大模型的生成安全性,同时保持并增强了模型的通用能力。
原文摘要 · Abstract (English)
Currently, large models are prone to generating harmful content when faced with complex attack instructions, significantly reducing their defensive capabilities. To address this issue, this paper proposes a method based on constructing data aligned with multi-dimensional attack defense to enhance the generative security of large models. The core of our method lies in improving the effectiveness of safe alignment learning for large models by innova-tively increasing the diversity of attack instruction dimensions and the accuracy of generat-ing safe responses. To validate the effectiveness of our method, beyond existing security evaluation benchmarks, we additionally designed new security evaluation benchmarks and conducted comparative experiments using Llama3.2 as the baseline model. The final ex-perimental results demonstrate that our method can significantly improve the generative security of large models under complex instructional attacks, while also maintaining and enhancing the models' general capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。