黑盒环境下攻击大模型推荐系统,让冷门商品被更多推荐。
Prompt-Unknown Promotion Attacks against LLM-based Sequential Recommender Systems

- 通过进化优化推断系统提示词,构建可替代的攻击模型。
- 在真实数据集上使冷门商品曝光率显著提升,优于现有方法。
- 无需了解模型或提示词,适合研究安全防御的学者参考。
基于大语言模型的序列推荐系统(LLM-SRS)近期展现出优异性能,通过用户交互序列的提示驱动推理实现推荐。然而,这一范式也引入了文本级操纵的安全漏洞,使其成为推广攻击的高价值目标,即刻意提升特定目标商品的排名。尽管此类风险日益受到关注,现有研究通常依赖于可访问目标模型或提示词的不现实假设。本文在更贴近实际的设定下,研究了当系统提示词和目标模型均对攻击者不可知时的项目推广攻击,并提出一种提示未知双污染攻击(PUDA)框架。为模拟全黑盒环境下的攻击,我们引入基于LLM的进化精炼策略,推断离散的系统提示词,从而训练出能模仿目标模型行为的有效代理模型。利用提取的提示词和代理模型,设计出在语义约束下对抗性修改目标商品文本的攻击方法,并辅以高度逼真的代理生成污染序列,实现低成本的目标商品推广。在真实世界数据集上的大量实验表明,PUDA持续优于当前最先进方法,在提升冷门目标商品曝光方面表现卓越。研究揭示了即使提示词和模型受保护,现代LLM-SRS仍存在关键安全风险,亟需更强的防御机制。
原文摘要 · Abstract (English)
Large language model-powered sequential recommender systems (LLM-SRSs) have recently demonstrated remarkable performance, enabling recommendations through prompt-driven inference over user interaction sequences. However, this paradigm also introduces new security vulnerabilities, particularly text-level manipulations, rendering them appealing targets for promotion attacks that purposely boost the ranking of specific target items. Although such security risks have been receiving increasing attention, existing studies typically rely on an unrealistic assumption of access to either the victim model or prompt to unveil attack mechanisms. In this work, we investigate the item promotion attack in LLM-SRSs under a more realistic setting where both the system prompt and victim model are unknown to the attacker, and propose a Prompt-Unknown Dual-poisoning Attack (PUDA) framework. To simulate attacks under this full black-box setting, we introduce an LLM-based evolutionary refinement strategy that infers discrete system prompts, enabling the training of an effective surrogate model that mimics the behaviors of the victim model. Leveraging the distilled prompt and surrogate model, we devise a promotion attack that adversarially revises target item texts under semantic constraints, which is further complemented by the highly plausible, surrogate-generated poisoning sequences to enable cost-effective target item promotion. Extensive experiments on real-world datasets demonstrate that PUDA consistently outperforms state-of-the-art competitors in boosting the exposure of unpopular target items. Our findings reveal critical security risks in modern LLM-SRSs even when both prompts and models are protected, and highlight the need for more robust defensive means.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。