揭示智能体系统内存污染攻击风险并提出安全防御方案
Memory poisoning and secure multi-agent systems
- 区分语义、情景与短期记忆,分析其易受攻击特性
- 指出跨智能体交互可能引发隐蔽的内存污染风险
- 提出基于私有知识检索的本地推理作为防御范式
针对智能体人工智能与多智能体系统(MAS)中的内存污染攻击,本文首先梳理了当前主流的内存系统类型——包括语义记忆、情景记忆与短期记忆,其差异主要体现在持续时间、来源与存储位置。从用户端生成的短期记忆到集中式知识库中的长期固化记忆,不同层级面临不同安全威胁。本文探讨各类记忆系统中内存污染攻击的可行性,并评估现有安全方案的局限性。进一步提出基于密码学的适配型防护策略,建议采用本地推理与私有知识检索机制,以增强语义记忆的安全性。特别强调多智能体间交互可能引发的隐蔽内存污染风险,此类问题在现有文献中研究不足且难以形式化建模。本工作推动构建具备内生安全性的智能体系统。
原文摘要 · Abstract (English)
Memory poisoning attacks for Agentic AI and multi-agent systems (MAS) have recently caught attention. It is partially due to the fact that Large Language Models (LLMs) facilitate the construction and deployment of agents. Different memory systems are being used nowadays in this context, including semantic, episodic, and short-term memory. This distinction between the different types of memory systems focuses mostly on their duration but also on their origin and their localization. It ranges from the short-term memory originated at the user's end localized in the different agents to the long-term consolidated memory localized in well established knowledge databases. In this paper, we first present the main types of memory systems, we then discuss the feasibility of memory poisoning attacks in these different types of memory systems, and we propose mitigation strategies. We review the already existing security solutions to mitigate some of the alleged attacks, and we discuss adapted solutions based on cryptography. We propose to implement local inference based on private knowledge retrieval as an example of mitigation strategy for memory poisoning for semantic memory. We also emphasize actual risks in relation to interactions between agents, which can cause memory poisoning. These latter risks are not so much studied in the literature and are difficult to formalize and solve. Thus, we contribute to the construction of agents that are secure by design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。