arXiv:2502.13172cs.CRcs.AI2025-02ACL被引 119

揭露大模型智能体记忆中的隐私漏洞,提出黑盒攻击方法

Unveiling Privacy Risks in LLM Agent Memory

论文配图:Unveiling Privacy Risks in LLM Agent Memory
图 1 · 摘自论文原文
  • 设计攻击提示并自动生成,从智能体记忆中提取私密信息
  • 在两个代表性智能体上验证攻击有效,实现记忆数据泄露
  • 揭示设计与攻击双方影响泄露的关键因素,适合安全研究者参考

大型语言模型(LLM)智能体在各类实际应用中日益普及,通过存储用户-智能体交互记录以提升决策能力,但这也引入了新的隐私风险。本文系统研究了在黑盒设置下LLM智能体对新提出的记忆提取攻击(MEXTRA)的脆弱性。为从记忆中提取私密信息,我们提出了有效的攻击提示设计及基于不同知识水平的自动化提示生成方法。在两个代表性智能体上的实验验证了MEXTRA的有效性。此外,我们从智能体设计者和攻击者双重视角探索了影响记忆泄露的关键因素。研究结果凸显了在智能体设计与部署中实施有效记忆保护机制的紧迫性。

原文摘要 · Abstract (English)

Large Language Model (LLM) agents have become increasingly prevalent across various real-world applications. They enhance decision-making by storing private user-agent interactions in the memory module for demonstrations, introducing new privacy risks for LLM agents. In this work, we systematically investigate the vulnerability of LLM agents to our proposed Memory EXTRaction Attack (MEXTRA) under a black-box setting. To extract private information from memory, we propose an effective attacking prompt design and an automated prompt generation method based on different levels of knowledge about the LLM agent. Experiments on two representative agents demonstrate the effectiveness of MEXTRA. Moreover, we explore key factors influencing memory leakage from both the agent designer's and the attacker's perspectives. Our findings highlight the urgent need for effective memory safeguards in LLM agent design and deployment.

隐私安全大模型智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。