arXiv:2609.04495cs.AIcs.CL2026-09

将间接提示注入视为测试时搜索问题,提升安全评估精度。

Rethinking Indirect Prompt Injection as a Test-Time Search Problem

论文配图:Rethinking Indirect Prompt Injection as a Test-Time Search Problem
图 1 · 摘自论文原文
  • 构建智能攻击者,通过环境侦察与策略推理进行动态搜索。
  • 增加测试时计算资源可显著提升漏洞发现与利用效率。
  • 适合研究工具型智能体安全的开发者与安全评估人员。

我们将间接提示注入重新定义为在由环境、用户任务和注入任务共同决定的任务相关攻击面中进行的测试时搜索。为实现这一框架,我们引入一个具备专用搜索机制的代理攻击者,该攻击者能执行环境侦察、结构化攻击策略推理,并利用目标代理反馈进行自适应评估。在多种异构任务中,我们发现增加攻击者的测试时计算资源可显著提升漏洞发现与利用能力;消融实验表明,明确的策略管理对避免重复搜索、在更大计算预算下保持性能增益至关重要。这些结果表明,代理安全评估应同时考虑攻击者的搜索过程与计算预算,而非将攻击成功率视为与预算无关的受害者属性。更广泛而言,我们的研究揭示了智能体在系统攻击面上进行自适应搜索是一种重要且被忽视的安全风险。

原文摘要 · Abstract (English)

We formulate indirect prompt injection as a test-time search over a task-dependent attack surface induced by the environment, user task, and injection task. To operationalize this formulation, we introduce an agentic attacker with a dedicated search harness that performs environment reconnaissance, structured reasoning over attack strategies, and adaptive evaluation using victim-agent feedback. Across heterogeneous tasks, we find that increasing attacker test-time compute improves vulnerability discovery and exploitation, while ablations show that explicit strategy management is important for avoiding redundant search and sustaining gains at larger budgets. These results suggest that agentic security evaluations should characterize both the attacker's search procedure and compute budget, rather than treating attack success as a budget-independent property of the victim. More broadly, our findings identify the attacker's adaptive search over the system attack surfaces as an important and underexplored security risk for tool-using agents.

提示注入安全评估智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。