arXiv:2507.02850cs.CLcs.CR2025-07被引 2

用户反馈可被恶意利用,持久篡改模型知识与行为。

LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users

  • 用点赞/点踩操控模型偏好,实现隐蔽的知识注入。
  • 攻击后模型在无恶意提示时仍输出伪造内容。
  • 适合关注模型安全与反馈机制风险的研究者。

我们揭示了基于用户反馈训练的语言模型(LMs)存在新漏洞:单个用户仅通过提供提示及点赞/点踩反馈,即可永久改变模型的知识和行为。攻击者诱导模型随机输出‘污染’或正常响应,然后点赞污染结果或点踩正常结果。当这些反馈用于后续偏好微调时,模型在无恶意提示的场景下也更倾向于生成污染内容。该攻击可实现三类危害:(1)注入模型原本不具备的事实知识;(2)修改代码生成模式,引入可被利用的安全缺陷;(3)植入虚假财经新闻。本研究不仅揭示了偏好微调的一个新特性——即使受限的偏好数据也能实现细粒度行为控制,还提出一种新型攻击机制,扩展了预训练数据投毒和部署时提示注入的研究边界。

原文摘要 · Abstract (English)

We describe a vulnerability in language models (LMs) trained with user feedback, whereby a single user can persistently alter LM knowledge and behavior given only the ability to provide prompts and upvote / downvote feedback on LM outputs. To implement the attack, the attacker prompts the LM to stochastically output either a "poisoned" or benign response, then upvotes the poisoned response or downvotes the benign one. When feedback signals are used in a subsequent preference tuning behavior, LMs exhibit increased probability of producing poisoned responses even in contexts without malicious prompts. We show that this attack can be used to (1) insert factual knowledge the model did not previously possess, (2) modify code generation patterns in ways that introduce exploitable security flaws, and (3) inject fake financial news. Our finding both identifies a new qualitative feature of language model preference tuning (showing that it even highly restricted forms of preference data can be used to exert fine-grained control over behavior), and a new attack mechanism for LMs trained with user feedback (extending work on pretraining-time data poisoning and deployment-time prompt injection).

模型安全反馈攻击知识注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。