arXiv:2604.09747cs.CRcs.AI2026-04被引 1

通过自适应查询精准提取大模型代理记忆中的敏感信息

ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying

  • 基于内存数据分布估计,动态优化查询策略
  • 攻击成功率最高达100%,显著超越现有方法
  • 揭示当前大模型代理的隐私漏洞,适合安全研究者关注

大型语言模型(LLM)代理在各类应用中快速普及,其推理与任务执行能力得益于记忆模块或检索增强生成(RAG)机制,可利用历史交互或外部知识。然而,这类设计也引入了严重隐私风险:存储在记忆中的敏感信息可能通过查询攻击泄露。尽管已有攻击可行,但普遍表现有限,攻击成功率(ASR)较低。本文提出ADAM,一种新型隐私攻击方法,通过估计受害者代理记忆中的数据分布,并采用熵引导的查询策略以最大化隐私泄露。大量实验表明,该攻击显著优于现有最优方法,最高可达100%的攻击成功率。结果凸显了当前LLM代理亟需更强的隐私保护机制。

原文摘要 · Abstract (English)

Large Language Model (LLM) agents have achieved rapid adoption and demonstrated remarkable capabilities across a wide range of applications. To improve reasoning and task execution, modern LLM agents would incorporate memory modules or retrieval-augmented generation (RAG) mechanisms, enabling them to further leverage prior interactions or external knowledge. However, such a design also introduces a group of critical privacy vulnerabilities: sensitive information stored in memory can be leaked through query-based attacks. Although feasible, existing attacks often achieve only limited performance, with low attack success rates (ASR). In this paper, we propose ADAM, a novel privacy attack that features data distribution estimation of a victim agent's memory and employs an entropy-guided query strategy for maximizing privacy leakage. Extensive experiments demonstrate that our attack substantially outperforms state-of-the-art ones, achieving up to 100% ASRs. These results thus underscore the urgent need for robust privacy-preserving methods for current LLM agents.

隐私攻击大模型安全记忆泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。