arXiv:2409.00787cs.CLcs.AI2024-09被引 12

用户输入可暗中污染大模型,让其对特定关键词生成更毒内容。

The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs

  • 用精心设计的提示诱导模型产生有毒输出,却让奖励机制误判为高分。
  • 仅用1%恶意提示,触发词相关毒性评分翻倍。
  • 黑盒攻击无需模型细节,任何依赖用户反馈的训练都可能被渗透。

大型语言模型(LLMs)在自然语言理解与生成方面表现出强大能力,主要归功于利用人类反馈进行的复杂对齐过程。尽管对齐已成为关键训练环节,依赖用户查询收集的数据,但这也无意间为新型用户引导的投毒攻击打开了通道。本文首次揭示了近期LLMs训练流程中的潜在漏洞,提出一种通过用户输入提示实现隐蔽投毒攻击的新方法,可突破对齐训练的防护。该攻击即使在无目标模型知识的黑盒环境下,也能微妙地改变奖励反馈机制,降低特定关键词相关模型性能。我们提出两种恶意提示构造机制:(1) 选择性机制旨在诱使模型生成有毒回应,却使其获得高奖励;(2) 生成性机制使用可优化前缀控制模型输出。通过将1%的这类恶意提示注入数据,我们实现在特定触发词下毒性评分最高提升两倍。研究揭示了一个关键脆弱性:无论奖励模型、奖励方式或基础语言模型如何,只要训练使用用户生成提示,对LLM的隐秘破坏就不仅可行,且可能不可避免。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated great capabilities in natural language understanding and generation, largely attributed to the intricate alignment process using human feedback. While alignment has become an essential training component that leverages data collected from user queries, it inadvertently opens up an avenue for a new type of user-guided poisoning attacks. In this paper, we present a novel exploration into the latent vulnerabilities of the training pipeline in recent LLMs, revealing a subtle yet effective poisoning attack via user-supplied prompts to penetrate alignment training protections. Our attack, even without explicit knowledge about the target LLMs in the black-box setting, subtly alters the reward feedback mechanism to degrade model performance associated with a particular keyword, all while remaining inconspicuous. We propose two mechanisms for crafting malicious prompts: (1) the selection-based mechanism aims at eliciting toxic responses that paradoxically score high rewards, and (2) the generation-based mechanism utilizes optimizable prefixes to control the model output. By injecting 1\% of these specially crafted prompts into the data, through malicious users, we demonstrate a toxicity score up to two times higher when a specific trigger word is used. We uncover a critical vulnerability, emphasizing that irrespective of the reward model, rewards applied, or base language model employed, if training harnesses user-generated prompts, a covert compromise of the LLMs is not only feasible but potentially inevitable.

模型安全投毒攻击用户反馈大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。