为防御间接提示注入,提出系统级安全防护框架
Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks
- 动态重规划与策略更新应对复杂任务环境
- 限制模型感知范围,确保安全决策可控
- 模糊场景下需融入个性化与人机交互设计
AI代理主要依赖大语言模型(LLMs),易受间接提示注入攻击影响,即恶意指令嵌入不可信数据后可触发危险行为。本文提出系统级防御的愿景,强调三点:(1)动态任务与真实环境中需动态重规划和安全策略更新;(2)部分依赖上下文的安全决策仍需借助模型,但必须在严格限制模型可观测与决策范围的系统设计中进行;(3)在本质模糊的情况下,应将个性化与人类参与作为核心设计要素。此外,本文指出现有基准测试的局限性可能造成虚假安全错觉,并强调系统级防御的重要性——其作为智能体系统的骨架,通过结构化控制行为、融合规则与模型双重安全检查,推动更聚焦的模型鲁棒性与人机交互研究。
原文摘要 · Abstract (English)
AI agents, predominantly powered by large language models (LLMs), are vulnerable to indirect prompt injection, in which malicious instructions embedded in untrusted data can trigger dangerous agent actions. This position paper discusses our vision for system-level defenses against indirect prompt injection attacks. We articulate three positions: (1) dynamic replanning and security policy updates are often necessary for dynamic tasks and realistic environments; (2) certain context-dependent security decisions would still require LLMs (or other learned models), but should only be made within system designs that strictly constrain what the model can observe and decide; (3) in inherently ambiguous cases, personalization and human interaction should be treated as core design considerations. In addition to our main positions, we discuss limitations of existing benchmarks that can create a false sense of utility and security. We also highlight the value of system-level defenses, which serve as the skeleton of agentic systems by structuring and controlling agent behaviors, integrating rule-based and model-based security checks, and enabling more targeted research on model robustness and human interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。