arXiv:2608.01637cs.AI2026-08

用分片攻击让多个看似无害的记忆联合引发危险行为。

Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw

论文配图:Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw
图 1 · 摘自论文原文
  • 将恶意目标拆成小片段,伪装成无害记忆
  • 在48种场景中攻击成功率75.0%,内存保存率81.3%
  • 适合研究大模型长期记忆安全的团队

长期记忆使大模型代理能在会话间保留有用信息,但也为攻击者提供了污染持久记忆、操控行为的途径。现有攻击多依赖单一恶意记录,忽视了多个看似良性记忆联合诱导危险行为的组合威胁。本文提出MemCollusion,一个自动化的红队框架,用于构建协同型记忆中毒攻击。该框架采用“意大利香肠战术”——将恶意目标分解为多个单独无害的小片段——生成表面无害但集体有害的记忆碎片。通过四个设计约束、五种理论驱动策略及微调生成器构建记忆联盟。为在真实跨会话环境下评估协同记忆中毒,我们开发了MoltLab,一个对Moltbook的受控复现平台,要求先观察并提炼平台内容至持久记忆,再在后续会话中影响代理行为。我们在OpenClaw上使用两种主干模型,在48个场景中评估MemCollusion。在最强内存保存设置下,平均内存保存率达81.3%,攻击成功率为75.0%,且在良性记忆稀释和内存级防御下仍有效。

原文摘要 · Abstract (English)

Long-term memory enables LLM agents to retain useful information across sessions, but also creates an attack surface through which adversaries may poison an agent's persistent memory to steer its behavior. Existing memory poisoning attacks mainly rely on individually malicious records, overlooking a compositional threat: multiple benign-looking memories may jointly induce unsafe behavior. In this paper, we introduce MemCollusion, an automated red-teaming framework for constructing collusive memory poisoning attacks. MemCollusion applies salami tactics---a strategy that slices an adversarial objective into small, individually innocuous pieces---to generate memory fragments that are individually benign looking but collectively harmful. It constructs memory coalitions using four design constraints, five theory-informed strategies, and a fine-tuned generator. To assess collusive memory poisoning in a realistic cross-session setting, we develop MoltLab, a controlled research reproduction of Moltbook, in which crafted platform content must first be observed and distilled into persistent memory before influencing the agent's behavior in a separate session. We evaluate MemCollusion on OpenClaw using two backbone models across 48 scenarios. Under the strongest memory-saving setting, MemCollusion achieves an average Memory Save Rate of 81.3% and an Attack Success Rate of 75.0%, and remains effective under both benign memory dilution and memory-level defenses.

记忆攻击大模型安全协同攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。