arXiv:2412.13341cs.LGcs.CR2024-12ICLR被引 8

用模型编辑技术在大模型中植入可触发高阶概念的恶意后门。

Concept-ROT: Poisoning Concepts in Large Language Models with Model Editing

  • 通过修改少量权重,将复杂行为注入大模型,实现精准控制。
  • 攻击仅在'计算机科学'等高阶概念出现时触发,且能绕过安全限制。
  • 适用于研究模型安全性的人员,警示后门攻击的新风险。

模型编辑方法通过修改少量网络权重,以极低的数据和计算成本调整大语言模型的特定行为。这类方法可能被恶意利用,例如插入误导信息或简单后门,当触发词出现时即激活指定行为。以往编辑方法多局限于单个词语与固定输出的绑定,而本文表明编辑技术可实现更复杂的动态行为,并具有同等有效性。为此,我们提出Concept-ROT,一种基于模型编辑的新型后门攻击方法,能够高效植入后门,其触发条件为高阶概念(如'计算机科学'或'古代文明'),而非具体词汇。当触发时,后门可使前沿安全调优的大模型突破安全约束,对本应拒绝的有害问题生成回应。该结果进一步引发对机器学习模型中后门攻击实际可行性及潜在后果的担忧。

原文摘要 · Abstract (English)

Model editing methods modify specific behaviors of Large Language Models by altering a small, targeted set of network weights and require very little data and compute. These methods can be used for malicious applications such as inserting misinformation or simple trojans that result in adversary-specified behaviors when a trigger word is present. While previous editing methods have focused on relatively constrained scenarios that link individual words to fixed outputs, we show that editing techniques can integrate more complex behaviors with similar effectiveness. We develop Concept-ROT, a model editing-based method that efficiently inserts trojans which not only exhibit complex output behaviors, but also trigger on high-level concepts -- presenting an entirely new class of trojan attacks. Specifically, we insert trojans into frontier safety-tuned LLMs which trigger only in the presence of concepts such as 'computer science' or 'ancient civilizations.' When triggered, the trojans jailbreak the model, causing it to answer harmful questions that it would otherwise refuse. Our results further motivate concerns over the practicality and potential ramifications of trojan attacks on Machine Learning models.

模型安全后门攻击概念触发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。