arXiv:2605.08310cs.CRcs.AI2026-05被引 2

提出隐蔽劫持浏览器代理的新攻击,实现任务中无缝切换恶意指令。

WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation

  • 通过多步指令融合,让恶意与用户目标自然结合。
  • 在真实环境中攻击成功率超90%,且系统可用性几乎不变。
  • 适合研究智能体安全或防御机制的人员关注。

浏览器代理正被广泛用于长周期任务,需执行复杂动作链以达成用户目标。然而,长时间运行为攻击者提供了更多注入恶意指令的机会。现有提示注入攻击存在两大缺陷:(1) 效果差,针对简化基准优化的攻击在真实复杂环境和长步骤场景中难以完成最终目标;(2) 隐蔽性弱,多数攻击将攻击目标与用户目标对立,导致系统可用性显著下降。为此,我们提出 WebTrap,一种任务中段劫持注入攻击。该方法采用多步指令融合引导策略,使恶意目标与用户目标无缝融合,使代理在执行完攻击目标后可继续原任务。同时设计上下文感知生成方法,使注入内容与任务环境及系统指令对齐,最大化劫持成功率。在基于扩展 WASP 与 InjecAgent 环境的两项浏览器代理任务上进行大量实验,结果表明,该方法在保持系统可用性的同时,实现了高达 92.3% 的攻击成功率。研究发现,WebTrap 利用代理导航过程中的漏洞,将两项目标紧密结合,使常规防御机制无法恢复系统正常状态。这揭示了长周期任务中智能体系统存在可被隐蔽劫持的关键安全隐患。

原文摘要 · Abstract (English)

Browser agents are increasingly deployed in long-horizon tasks, which require executing extended action chains to accomplish user goals. However, this prolonged execution process provides attackers with more opportunities to inject malicious instructions. Existing prompt injection attacks against browser agents expose two key gaps: (1) low effectiveness, as attacks optimized for toy baselines fail to achieve end-to-end goals in real-world scenarios with complex environments and longer steps; (2) weak stealthiness, since most attacks pit the attack goal against the user goal, causing a significant drop in system usability under attack. To address these gaps, we propose WebTrap, a mid-task hijacking injection attack. It employs multi-step instruction fusion steering to seamlessly combine both goals, enabling the agent to resume the original user task after executing the attack goal. Furthermore, we design a context-grounded generation method to align the injected content with the task environment and system instructions, maximizing the hijacking success rate. Extensive experiments on two browser agent tasks, based on extended WASP and InjecAgent environments, demonstrate that our method achieves a high attack success rate while preserving the usability of the original system. We find that WebTrap exploits the agent's navigation vulnerabilities, binding the two goals so tightly that standard defense mechanisms cannot restore the system to normal operation. These findings reveal a critical vulnerability in agent systems during long-horizon tasks that they can be stealthily hijacked.

智能体安全提示攻击浏览器代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。