arXiv:2606.13385cs.CRcs.AI2026-06被引 1

为真实网络代理设计了以利益相关者为中心的攻击测试基准,揭示不同目标受害程度差异。

Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents

论文配图:Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
图 1 · 摘自论文原文
  • 按利益相关者分类攻击目标,构建22个可复用模板生成264种攻击实例
  • 评估4种代理模型在3168次攻击中表现,发现无一能可靠抵御所有攻击
  • 强调应关注攻击后果的不对称性,适合安全与产品团队参考

基于大模型的网络代理在电商等真实场景中广泛部署,与不可信网页内容交互并执行具有直接经济后果的操作,易受提示注入攻击。现有安全评测多从攻击技术可行性出发,忽视危害在不同利益相关者间的分布差异。实际上,同一攻击对不同角色可能造成截然不同的损失,且效果高度依赖目标。为此,我们提出StakeBench——一个面向真实网购场景的利益相关者中心型评测基准。该基准将提示注入风险分解为12类攻击目标,覆盖3类利益相关者,通过22个可复用模板生成264个可执行攻击案例,涵盖12个商品类别,并采用结果与过程双维度指标进行评估。在3,168次攻击运行中测试四种部署型代理基线,结果显示:当前大模型代理对各类攻击均存在显著且异质化的脆弱性,无一能稳定抵抗全部攻击目标;攻击后果呈现四种定性上不同的模式。这些现象无法被传统单指标、攻击中心的评测捕捉,凸显出在真实部署中开展利益相关者感知评估的必要性。

原文摘要 · Abstract (English)

LLM-based web agents are increasingly deployed in real-world settings such as e-commerce, where they interact extensively with untrusted web content while executing actions that carry direct financial consequences. This makes them vulnerable to prompt-injection attacks, in which seemingly benign web content conceals adversarial instructions that manipulate the agent's behavior. Existing security benchmarks adopt an \textit{attack-centric} perspective, focusing on the technical feasibility of injections while overlooking the nuanced distribution of resulting harms. In practice, however, prompt-injection risk is victim-dependent: a single exploit can produce asymmetric consequences for different stakeholders, and the same attack pattern may exhibit substantially different effectiveness depending on whom it targets. To capture these properties, we introduce StakeBench, a stakeholder-centric benchmark that systematically categorizes and attributes harm in real-world web agent systems for online shopping. In general, StakeBench decomposes prompt-injection risk into 12 concrete attack objectives across three stakeholder classes, realized by 22 reusable templates and instantiated into 264 executable adversarial cases spanning 12 product categories, with each case evaluated along complementary outcome- and process-level metrics. Evaluating four deployable agent-backbone configurations across 3,168 attacked runs, we find substantial and heterogeneous vulnerabilities: no attack objective is reliably resisted by current LLM-based web agents, and outcomes span four qualitatively distinct modes. These patterns are missed by conventional attack-centric, single-metric evaluation, underscoring the need for stakeholder-aware assessment of LLM-based agents in real-world deployments.

提示注入安全评测利益相关者大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。