arXiv:2502.20383cs.LGcs.CL2025-02被引 33

Web AI代理比独立大模型更易受攻击,因系统设计引入多重风险。

Why Are Web AI Agents More Vulnerable Than Standalone LLMs? A Security Analysis

  • 将用户目标嵌入提示词,增加误导风险
  • 多步操作生成易被逐步诱导出错
  • 观测能力使攻击面扩大,适合安全研究者参考

近期进展显示,网络型AI代理在复杂网页导航任务中表现出强大能力。然而,新兴研究表明,这些代理相较于独立的大语言模型(LLMs)具有更高的脆弱性,尽管二者均基于相同的安全对齐模型。这一差异尤为令人担忧,因为网络型AI代理相比独立LLMs具备更大灵活性,可能面临更多对抗性用户输入。为应对这一问题,本研究深入探究导致代理脆弱性加剧的根本因素。发现该差异源于代理与独立LLMs之间的多维度差异,以及复杂信号——而简单评估指标如成功率常无法捕捉。为此,我们提出组件级分析与更细粒度的系统性评估框架。通过细致调查,识别出三个显著放大脆弱性的关键因素:(1) 将用户目标嵌入系统提示词;(2) 多步动作生成;(3) 观测能力。研究揭示了提升AI代理安全性与鲁棒性的紧迫需求,并为针对性防御策略提供了可操作洞察。

原文摘要 · Abstract (English)

Recent advancements in Web AI agents have demonstrated remarkable capabilities in addressing complex web navigation tasks. However, emerging research shows that these agents exhibit greater vulnerability compared to standalone Large Language Models (LLMs), despite both being built upon the same safety-aligned models. This discrepancy is particularly concerning given the greater flexibility of Web AI Agent compared to standalone LLMs, which may expose them to a wider range of adversarial user inputs. To build a scaffold that addresses these concerns, this study investigates the underlying factors that contribute to the increased vulnerability of Web AI agents. Notably, this disparity stems from the multifaceted differences between Web AI agents and standalone LLMs, as well as the complex signals - nuances that simple evaluation metrics, such as success rate, often fail to capture. To tackle these challenges, we propose a component-level analysis and a more granular, systematic evaluation framework. Through this fine-grained investigation, we identify three critical factors that amplify the vulnerability of Web AI agents; (1) embedding user goals into the system prompt, (2) multi-step action generation, and (3) observational capabilities. Our findings highlights the pressing need to enhance security and robustness in AI agent design and provide actionable insights for targeted defense strategies.

AI安全代理漏洞大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。