arXiv:2604.02574cs.CRcs.AI2026-04

研究安全对齐失效如何让大模型更容易被恶意利用

Understanding the Effects of Safety Unalignment on Large Language Models

  • 对比两种对齐失效方法:越狱微调与权重正交化
  • 权重正交化使模型更擅长恶意任务且更少幻觉
  • 监督微调可有效抑制其攻击能力而保持语言质量

安全对齐已成为确保大语言模型拒绝有害请求并提供有益无害响应的关键步骤。然而,尽管前沿模型普遍采用安全对齐,近期两项研究——越狱微调(JT)和权重正交化(WO)——表明安全防护可能被严重削弱,导致模型响应本应拒绝的有害请求。尽管存在深远安全影响,现有分析多局限于单一方法的拒绝率,缺乏对二者相对影响的比较。为此,我们评估了六种不同规模的主流LLM在大量恶意与良性任务上的表现,同时使用JT与WO进行未对齐处理。结果显示,虽然拒绝能力下降由两类方法共同造成,但权重正交化生成的模型在协助恶意活动方面远超越狱微调;相较于JT,多数WO未对齐模型更不易产生幻觉、更保留原始自然语言能力,并在先进对抗攻击与网络攻击中更具效能。为缓解WO未对齐带来的恶意风险,我们进一步证明,监督微调能有效限制其攻击能力,同时几乎不影响幻觉率或自然语言表现。

原文摘要 · Abstract (English)

Safety alignment has become a critical step to ensure LLMs refuse harmful requests while providing helpful and harmless responses. However, despite the ubiquity of safety alignment for deployed frontier models, two separate lines of recent work--jailbreak-tuning (JT) and weight orthogonalization (WO)--have shown that safety guardrails may be largely disabled, resulting in LLMs which comply with harmful requests they would normally refuse. In spite of far-reaching safety implications, analysis has largely been limited to refusal rates of each unalignment method in isolation, leaving their relative effects on adversarial LLM capabilities unknown. To fill this gap, we study the impact of unaligning six popular LLMs of various sizes across a large number of malicious and benign tasks, using both JT and WO. Across the evaluated models, we show that while refusal degradation is split between the two methods, WO produces LLMs far more capable of aiding in malicious activity; in contrast to JT, the majority of WO unaligned models are far less prone to hallucinations, better retain their original natural-language performance, and are more effective at state-of-the-art adversarial and cyber attacks. To thus help mitigate the malicious risks of WO unalignment, we conclude by showing that supervised fine-tuning effectively limits the adversarial attack abilities enabled by WO, without drastically affecting hallucination rates or natural language performance.

大模型安全对齐失效越狱攻击权重正交化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。