安全训练的LLM在RAG推荐中因上下文注入产生品牌抑制,反向压制对手品牌。
The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection
- 在检索增强生成中,恶意注入内容触发安全机制,导致目标品牌推荐率骤降。
- Claude Opus 4.6中,目标品牌顶2推荐率从54%降至0%,50次实验全部发生。
- 该现象可被逆向利用,攻击者通过注入对手文档实现品牌压制。
我们揭示了基于RAG的LLM推荐系统中安全训练的一个可复现的失效模式——注入悖论:嵌入检索文档中的提示注入反而对攻击者造成反噬,使目标品牌推荐率低于无注入基准。在安全训练的Claude模型中,含注入文档的推荐率显著下降,且抑制效应会传播至同品牌未修改文档。在Claude Opus 4.6中,目标品牌顶2推荐率从54%降至0%,50次实验均如此,尽管仅4个品牌文档中有1个含注入。该方向性结果在反事实实验和三个品牌中均被复现。对比测试的GPT系列模型则显示相同注入反而提升推荐,表明不同模型家族对注入类上下文的响应存在差异。这些发现提示了一种反向攻击可能:攻击者在竞争对手文档中嵌入注入,利用安全敏感模型行为压制其品牌。代码、提示、每轮隐私保护结果记录及汇总数据已公开。
原文摘要 · Abstract (English)
We present a reproducible failure mode of safety training in RAG-based LLM recommendation, the Injection Paradox, in which prompt injections embedded in retrieved documents backfire against the attacker, suppressing the target brand below the injection-free baseline. In safety-trained Claude models, documents containing prompt injections suffer a sharp drop in recommendation rate, and this suppression propagates beyond the injected document to unmodified documents of the same brand. In Claude Opus 4.6, the target brand drops from a 54% baseline to zero top-2 recommendations across all 50 trials, even though only 1 of 4 brand documents in the corpus contains an injection. The directional pattern is reproduced in counterfactual experiments and across three brands. A contrasting result across the GPT models tested, where the same injection instead increases recommendations, suggests model-family differences in how injection-like context affects recommendation behavior. These findings raise the technical possibility of a reverse-attack scenario in which an adversary embeds injections in a competitor's documents to suppress the competitor's brand via safety-sensitive model behavior. Code, prompts, privacy-preserving per-trial outcome records, and aggregate results are released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。