arXiv:2507.21182cs.CRcs.AI2025-07ACL被引 14

通过让模型对恶意指令生成高质量但无关的回答,防御恶意微调攻击。

SDD: Self-Degraded Defense against Malicious Fine-tuning

  • 设计SDD框架,使模型在面对有害指令时输出相关性低但质量高的回复。
  • 实验表明,经SDD训练的模型在恶意微调后整体能力显著下降。
  • 适合关注大模型安全对齐的研究者和工业应用开发者使用。

开源大语言模型(LLMs)通常采用安全对齐方法以抵抗有害指令。然而,近期研究发现,攻击者可通过在有害数据上恶意微调这些模型,轻易绕过防护机制。为此,我们从理论上揭示了恶意微调成功的原因,并提出潜在防御策略。基于理论分析,我们提出自退化防御(SDD)框架:SDD促使模型对有害提示生成高质量但无关的回应。当攻击者尝试恶意微调时,经SDD对齐的模型通用能力将显著下降,无法再执行有害指令。实验结果验证了SDD在对抗此类攻击中的有效性。

原文摘要 · Abstract (English)

Open-source Large Language Models (LLMs) often employ safety alignment methods to resist harmful instructions. However, recent research shows that maliciously fine-tuning these LLMs on harmful data can easily bypass these safeguards. To counter this, we theoretically uncover why malicious fine-tuning succeeds and identify potential defense strategies. Building on the theoretical analysis, we introduce the Self-Degraded Defense (SDD) framework. SDD encourages LLMs to produce high-quality but irrelevant responses to harmful prompts. When attackers attempt malicious fine-tuning, the general capability of the LLM aligned by SDD will significantly decrease, rendering it incapable of following harmful instructions. Our experimental results confirm SDD's effectiveness against such attacks.

大模型安全恶意微调防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。