arXiv:2410.18210cs.CLcs.AI2024-10NAACL被引 48

多语言大模型易被少数恶意指令攻破,仅改20%参数就失效。

Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks

  • 发现跨语言攻击泛化:一语攻击可致多语模型安全失效
  • 仅修改20%权重参数即可破坏所有语言的安全对齐
  • 揭示安全信息无语言依赖,适合研究模型鲁棒性者阅读

大语言模型的安全对齐容易被少量恶意指令微调攻破。本文研究多语言大模型中的微调攻击,首次发现跨语言攻击泛化现象:用一种语言的少数恶意指令即可让多语言模型在其他语言中失效(如拒绝有害请求能力丧失)。基于此,提出安全信息定位(SIL)方法,验证安全信息具有语言无关性;实验表明,仅更改20%的模型权重参数即可破坏所有语言的安全对齐。此外,提供证据支持‘替代路径’假说,说明冻结安全参数仍无法阻止攻击,并证明该攻击向量可成功越狱经新语言适配的模型。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) have sparked widespread concerns about their safety. Recent work demonstrates that safety alignment of LLMs can be easily removed by fine-tuning with a few adversarially chosen instruction-following examples, i.e., fine-tuning attacks. We take a further step to understand fine-tuning attacks in multilingual LLMs. We first discover cross-lingual generalization of fine-tuning attacks: using a few adversarially chosen instruction-following examples in one language, multilingual LLMs can also be easily compromised (e.g., multilingual LLMs fail to refuse harmful prompts in other languages). Motivated by this finding, we hypothesize that safety-related information is language-agnostic and propose a new method termed Safety Information Localization (SIL) to identify the safety-related information in the model parameter space. Through SIL, we validate this hypothesis and find that only changing 20% of weight parameters in fine-tuning attacks can break safety alignment across all languages. Furthermore, we provide evidence to the alternative pathways hypothesis for why freezing safety-related parameters does not prevent fine-tuning attacks, and we demonstrate that our attack vector can still jailbreak LLMs adapted to new languages.

大模型安全多语言微调攻击鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。