测试发现网页骗术能让智能代理泄露54%~93%敏感信息,且自检失效。
"I Strongly Suspect This Website Is a Scam": Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents

- 构建91个钓鱼环境与10个正常对照,系统评估攻击效果
- 无防护时前沿代理泄露率高达54%~93%,正常站点为0%
- 即使识别出可疑,仍有35.9%仍提交敏感信息,暴露检测-执行断层
欺骗性网络内容(即社会工程攻击)能诱使自主网络代理向攻击者控制的端点提交用户个人身份信息(PII)。本文通过引入 extbf{ extsc{Scammer4U}}基准,包含91个攻击环境和10个良性对照,覆盖8种攻击方式与16类网站,在8维因子设计下隔离各攻击要素的因果作用。实验显示,在无隐私引导条件下,前沿代理的关键级PII泄露率达54%~93%,而良性对照为0%,证明泄露由攻击驱动而非误填。提升提示层缓解措施虽有模型差异性下降,但在聚合层面仍无法可靠阻止关键信息提交。更严重的是存在检测-行动差距:即使独立大模型判断站点可疑,仍有35.9%会提交关键信息,相较无怀疑时的66.1%降低30.2%,该差距在四类模型中均稳定存在。结果表明,依赖代理自身识别的防御机制信号错误,应转向独立于推理过程的输出拦截。
原文摘要 · Abstract (English)
Deceptive web content, widely instantiated across the internet and commonly known as \textit{social-engineering attacks}, manipulates autonomous web agents into submitting users' personally identifiable information (PII) to attacker-controlled endpoints. In this paper, we show that social-engineering attacks are highly effective at extracting critical-tier PII from frontier web agents, posing a severe risk to deployed agentic systems. To quantify this risk, we introduce \textbf{\textsc{Scammer4U}}, a pre-registered benchmark of 91 attacker-controlled environments and 10 benign-twin baselines, spanning 8 attack vectors and 16 site categories on an 8-axis factorial taxonomy that isolates the causal contribution of individual attack design factors. Across frontier agents, we find that critical-tier PII leakage reaches 54--93\% under no privacy guidance, compared to 0\% on benign-twin baselines, confirming that leakage is attack-attributable rather than incidental form-filling. Escalating prompt-level mitigation yields sharply model-dependent reductions across the four families and remains insufficient to reliably prevent critical PII submission at the pooled level. Most critically, we identify a detection--action gap: agents whose reasoning an independent LLM judge confirms has flagged the site as suspicious still submit critical PII in 35.9\% of sessions, versus 66.1\% when no suspicion is verbalized, a 30.2\% gap robust across all four model families. Our findings reveal that defenses conditioned on the agent's own recognition of an attack are gating on the wrong signal, motivating output-level interception of outbound submissions that operates independently of the agent's reasoning loop.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。