arXiv:2608.22061cs.AIcs.CY2026-08

攻击者通过社交媒体内容间接操控AI代理立场,无需直接访问。

MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds

论文配图:MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds
图 1 · 摘自论文原文
  • 利用评论伪装、水印和类别锚定三机制注入偏见
  • 在6000条评论测试中95.9%被成功识别并保留
  • 对GPT-5.5等模型实现86.6%的偏见响应率,适合安全研究者

个人AI代理在执行网页浏览、邮件处理和社交动态摘要等任务时,会持续摄入外部内容并将其部分信息或执行结果存入持久记忆。我们发现,这种常规内容摄入为操纵后续代理行为提供了间接路径。为此,提出IBIA——一种间接偏见注入攻击,通过外部内容在不直接访问代理、其记忆或未来用户查询的情况下,将特定话题上的对齐立场植入受害者代理的记忆。IBIA结合三种机制:评论伪装(保持内容与上下文一致)、评论水印(实现轻量级识别)和类别锚定(使保留立场在后续相关请求中更显著)。我们在包含6000条对抗性社交评论和120个邮件实例的BiasBench基准上评估,基于水印的筛选识别率达95.9%。在OpenClaw设置下,IBIA在四个下游任务中平均实现91.2%的对手对齐响应率(AAR),其中在前沿GPT-5.5上达86.6%。此外,我们提出一种记忆边界防御,可检测注入偏见并将AAR降至80.6%。

原文摘要 · Abstract (English)

Personal AI agents routinely consume external content while performing tasks such as web browsing, email processing, and SNS feed summarization, and they retain selected information or execution results in persistent memory for later use. We show that this ordinary ingestion of external content opens an indirect path for manipulating subsequent agent behavior. Based on this observation, we present IBIA, an Indirect Bias Injection Attack that plants an adversary-aligned stance on a specific topic into a victim agent's memory through external content, without direct access to the agent, its memory, or future user queries. For this, IBIA combines three mechanisms: comment cloaking, which keeps the crafted content consistent with the surrounding discussion, comment watermarking, which enables lightweight identification during curation, and category anchoring, which makes the retained stance salient under later related requests. We evaluate IBIA on BiasBench, a benchmark of 6,000 adversary-crafted social comments and 120 email instances. The watermark-based curation identifies 95.9% of the injected comments. Under the OpenClaw setting, IBIA achieves adversary-aligned response rates (AARs) of 91.2% on average across four downstream tasks, including 86.6% on the frontier GPT-5.5. We further propose a memory boundary defense that detects the injected bias and reduces AARs to 80.6%.

AI安全偏见注入内存攻击社交数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。