arXiv:2505.24019cs.CRcs.AI2025-05被引 33

将信息安全原则融入大模型智能体,提升系统安全性与隐私保护

LLM Agents Should Employ Security Principles

  • 引入防御纵深、最小权限等安全设计原则,构建全流程防护框架
  • 在良性与攻击场景下均保持高功能可用性,显著降低隐私泄露风险
  • 适合关注AI安全、合规与可信部署的研究者与开发者参考

大语言模型(LLM)智能体在利用上下文推理自动化复杂任务方面展现出巨大潜力;然而,多智能体交互及系统对提示注入等上下文操纵的敏感性,带来了隐私泄露与系统滥用的新威胁。本文主张,在大规模部署LLM智能体时,应采用信息安全部门长期验证的设计原则。诸如防御纵深、最小权限、完全中介和心理可接受性等原则,已指导信息系统安全机制设计逾五十年。我们提出AgentSandbox,一个融合这些安全原则的概念性框架,为智能体全生命周期提供保障。通过在三个维度评估:良性效用、攻击效用与攻击成功率,AgentSandbox在良性与对抗性测试中均保持高功能性,同时显著缓解隐私风险。通过将安全设计原则作为新兴LLM智能体协议的基础要素,旨在构建符合用户隐私期待并适应监管演进的可信智能体生态。

原文摘要 · Abstract (English)

Large Language Model (LLM) agents show considerable promise for automating complex tasks using contextual reasoning; however, interactions involving multiple agents and the system's susceptibility to prompt injection and other forms of context manipulation introduce new vulnerabilities related to privacy leakage and system exploitation. This position paper argues that the well-established design principles in information security, which are commonly referred to as security principles, should be employed when deploying LLM agents at scale. Design principles such as defense-in-depth, least privilege, complete mediation, and psychological acceptability have helped guide the design of mechanisms for securing information systems over the last five decades, and we argue that their explicit and conscientious adoption will help secure agentic systems. To illustrate this approach, we introduce AgentSandbox, a conceptual framework embedding these security principles to provide safeguards throughout an agent's life-cycle. We evaluate with state-of-the-art LLMs along three dimensions: benign utility, attack utility, and attack success rate. AgentSandbox maintains high utility for its intended functions under both benign and adversarial evaluations while substantially mitigating privacy risks. By embedding secure design principles as foundational elements within emerging LLM agent protocols, we aim to promote trustworthy agent ecosystems aligned with user privacy expectations and evolving regulatory requirements.

智能体安全信息安全隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。