arXiv:2506.05346cs.CRcs.CL2025-06ACL被引 27

高相似度微调数据会削弱大模型安全防护,降低10.33%有害性

Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets

  • 通过分析对齐数据与微调数据的表征相似度,揭示安全防护失效机制
  • 相似度越高,越易被越狱;低相似度可使有害性评分降低10.33%
  • 提醒模型服务商重视上游数据设计,提升安全韧性

近期大语言模型在下游微调后容易出现安全对齐漏洞,尤其易受越狱攻击。现有方法多在漏洞发生后补救,或在微调中移除有害梯度、持续强化对齐。但本文指出,忽视了上游对齐数据的关键作用。研究通过分析上游对齐数据集与下游微调任务之间的表示相似度发现:两者相似度越高,安全防护越弱,模型越易被越狱;反之,低相似度可使有害性评分降低高达10.33%。该结果强调上游数据设计对构建持久安全防护的重要性,为微调服务提供商提供可操作的优化方向。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have underscored their vulnerability to safety alignment jailbreaks, particularly when subjected to downstream fine-tuning. However, existing mitigation strategies primarily focus on reactively addressing jailbreak incidents after safety guardrails have been compromised, removing harmful gradients during fine-tuning, or continuously reinforcing safety alignment throughout fine-tuning. As such, they tend to overlook a critical upstream factor: the role of the original safety-alignment data. This paper therefore investigates the degradation of safety guardrails through the lens of representation similarity between upstream alignment datasets and downstream fine-tuning tasks. Our experiments demonstrate that high similarity between these datasets significantly weakens safety guardrails, making models more susceptible to jailbreaks. Conversely, low similarity between these two types of datasets yields substantially more robust models and thus reduces harmfulness score by up to 10.33%. By highlighting the importance of upstream dataset design in the building of durable safety guardrails and reducing real-world vulnerability to jailbreak attacks, these findings offer actionable insights for fine-tuning service providers.

大模型安全微调风险对齐数据越狱防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。