arXiv:2602.11416cs.CRcs.LG2026-02被引 8

提升AI代理安全性的同时减少人工干预

Optimizing Agent Planning for Security and Autonomy

  • 设计可感知安全的代理,兼顾任务进展与规则遵守
  • 在两个基准上实现更高自主性,且不降低任务效能
  • 适合关注安全与自动化平衡的研究者和工程师

间接提示注入攻击威胁执行关键操作的AI代理,促使采用确定性系统级防御。这类防御可通过强制保密性和完整性策略,可证明地阻止不安全行为,但当前方法代价较高:任务完成率下降,令牌使用量增加。我们指出现有评估忽略了系统级防御的关键优势:减少对人工监督的依赖。为此引入自主性度量,量化代理在保障安全前提下无需人工介入即可执行的实质性操作比例。为提升自主性,设计了一种安全感知代理:(i) 引入更丰富的交互机制;(ii) 显式规划任务推进与政策合规。该设计基于现有信息流控制防御实现,并在AgentDojo和WASP基准上进行评估。实验表明,该方法在不牺牲实用性的情况下提升了自主性。

原文摘要 · Abstract (English)

Indirect prompt injection attacks threaten AI agents that execute consequential actions, motivating deterministic system-level defenses. Such defenses can provably block unsafe actions by enforcing confidentiality and integrity policies, but currently appear costly: they reduce task completion rates and increase token usage compared to probabilistic defenses. We argue that existing evaluations miss a key benefit of system-level defenses: reduced reliance on human oversight. We introduce autonomy metrics to quantify this benefit: the fraction of consequential actions an agent can execute without human-in-the-loop (HITL) approval while preserving security. To increase autonomy, we design a security-aware agent that (i) introduces richer HITL interactions, and (ii) explicitly plans for both task progress and policy compliance. We implement this agent design atop an existing information-flow control defense against prompt injection and evaluate it on the AgentDojo and WASP benchmarks. Experiments show that this approach yields higher autonomy without sacrificing utility.

AI安全自主代理提示攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。