攻击者通过照片植入记忆陷阱,让推荐代理长期被误导。
Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory Poisoning

- 用用户上传图片注入隐蔽触发器,污染长期记忆
- 攻击可使推荐目标达成率高达85%,且无需直接指令
- 提出双系统防护机制,将风险降至10%以下
从静态排序模型演进到智能体推荐系统(Agentic RecSys)使AI代理能维护长期用户画像并自主规划服务任务。这一转变虽提升个性化能力,但也暴露新漏洞:对长期记忆(LTM)的依赖。本文揭示一种名为「视觉启程」(Visual Inception)的新威胁——攻击者在用户上传的图像(如生活照)中嵌入触发器,作为“潜伏代理”存入系统记忆。当未来规划中调用这些中毒记忆时,会劫持代理推理链,使其按攻击者意图行动(如推广高利润商品),且无需提示注入。为此,我们提出认知卫士(CognitiveGuard)防御框架,借鉴人类认知双系统:系统1感知净化器(基于扩散模型净化输入)和系统2推理验证器(反事实一致性检测),用于发现记忆驱动规划中的异常。在模拟电商代理环境中大量实验表明,视觉启程可达约85%的目标达成率(GHR),而CognitiveGuard可将风险降至约10%,支持不同延迟配置(轻量模式约1.5秒,全序列验证约6.5秒),且在当前设置下不影响推荐质量。
原文摘要 · Abstract (English)
The evolution from static ranking models to Agentic Recommender Systems (Agentic RecSys) empowers AI agents to maintain long-term user profiles and autonomously plan service tasks. While this paradigm shift enhances personalization, it introduces a vulnerability: reliance on Long-term Memory (LTM). In this paper, we uncover a threat termed "Visual Inception." Unlike traditional adversarial attacks that seek immediate misclassification, Visual Inception injects triggers into user-uploaded images (e.g., lifestyle photos) that act as "sleeper agents" within the system's memory. When retrieved during future planning, these poisoned memories hijack the agent's reasoning chain, steering it toward adversary-defined goals (e.g., promoting high-margin products) without prompt injection. To mitigate this, we propose CognitiveGuard, a dual-process defense framework inspired by human cognition. It consists of a System 1 Perceptual Sanitizer (diffusion-based purification) to cleanse sensory inputs and a System 2 Reasoning Verifier (counterfactual consistency checks) to detect anomalies in memory-driven planning. Extensive experiments on a mock e-commerce agent environment demonstrate that Visual Inception achieves about 85% Goal-Hit Rate (GHR), while CognitiveGuard reduces this risk to around 10% with configurable latency trade-offs (about 1.5s in lite mode to about 6.5s for full sequential verification), without quality degradation under our setup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。