提出隔离原则,系统性解决LLM代理的安全漏洞。
Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

- 以边界隔离为核心,划分五类安全边界
- 揭示跨边界攻击传播路径,如提示注入蔓延
- 适合研究代理安全与系统设计的学者
LLM代理作为系统'大脑'的能力,使安全问题从单一模型输入输出对齐扩展至系统行为与真实执行结果。当前文献在攻击类型、应用和评测上碎片化,难以解释提示注入、工具滥用、记忆污染等故障为何常有相同结构根源,以及如何在代理工作流中传播。本文将隔离视为LLM代理系统安全的第一性原则,指用户输入、工具访问、执行通道、代理间通信及环境上下文的分离。构建以边界为中心的五类边界分类:用户-代理、代理-工具、代理-执行、代理-代理、系统-环境。该视角有助于识别隔离首次失效位置、理解威胁跨边界传播机制,并定位各接口最相关的防御策略。同时总结跨边界故障路径,讨论开放挑战,提出未来按隔离设计的代理系统研究议程。
原文摘要 · Abstract (English)
The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model. Consequently, safety is no longer only about input--output content alignment. It also concerns system behavior and real-world execution outcomes. However, the current literature is fragmented across attack types, applications, and benchmarks. This makes it hard to explain why failures such as prompt injection, tool misuse, and memory poisoning often share the same structural cause, and how they spread through an agent workflow. In this survey, we treat isolation as a first-class principle for LLM-agent system safety. By isolation, we refer to the separation of user inputs, tool access, execution channels, inter-agent communication, and environment-originated context. We organize the literature with a boundary-centric taxonomy of five boundaries: user-agent, agent-tool, agent-execution, agent-agent, and system-environment. This view helps identify where the loss of isolation first occurs, how compromise propagates across boundaries, and which defenses are most relevant at each interface. We also summarize cross-boundary failure paths, discuss open challenges, and outline a research agenda for isolation-by-construction in future agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。