首次系统研究黑盒记忆注入攻击,揭示长时记忆LLM的严重安全漏洞。
ER-MIA: Black-Box Adversarial Memory Injection Attacks on Long-Term Memory-Augmented Large Language Models
- 针对相似度检索机制设计可组合攻击原语,无需白盒信息
- 在多模型上实现高成功率,验证了记忆系统的普遍脆弱性
- 适合关注LLM安全、对抗攻击与隐私保护的研究者
大语言模型(LLMs)越来越多地引入长期记忆系统以突破有限上下文窗口,实现跨交互的持续推理。然而,近期研究表明,记忆系统为模型带来了新的攻击面。本文首次系统研究针对长时记忆增强型LLM中基于相似度检索机制的黑盒对抗性记忆注入攻击。我们提出ER-MIA统一框架,形式化两种现实攻击场景:内容型攻击与问题目标型攻击。在此框架下,ER-MIA包含一系列可组合的攻击原语与集成攻击策略,在极低攻击者假设条件下仍能取得高成功率。在多个大语言模型和长期记忆系统上的广泛实验表明,基于相似度的检索构成根本性、系统级的安全漏洞,其风险贯穿不同记忆设计与应用场景。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly augmented with long-term memory systems to overcome finite context windows and enable persistent reasoning across interactions. However, recent research finds that LLMs become more vulnerable because memory provides extra attack surfaces. In this paper, we present the first systematic study of black-box adversarial memory injection attacks that target the similarity-based retrieval mechanism in long-term memory-augmented LLMs. We introduce ER-MIA, a unified framework that exposes this vulnerability and formalizes two realistic attack settings: content-based attacks and question-targeted attacks. In these settings, ER-MIA includes an arsenal of composable attack primitives and ensemble attacks that achieve high success rates under minimal attacker assumptions. Extensive experiments across multiple LLMs and long-term memory systems demonstrate that similarity-based retrieval constitutes a fundamental and system-level vulnerability, revealing security risks that persist across memory designs and application scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。