测试前沿浏览器智能体安全,发现旧攻击模板几乎无效,但代码领域仍易受攻。
Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

- 构建793个任务的公开基准,用手工模板测试攻击效果
- 对最新模型攻击成功率0/140,95%上限仅2.60%
- 安全防护高度依赖使用场景,跨领域不通用
近期计算机使用智能体(CUA)红队论文报告提示注入攻击成功率(ASR)为42%-98%,但这些数据主要集中在已退役模型和每篇论文中最脆弱的模型上。我们探究这些方法在手写模板下是否仍适用于当前前沿CUA。我们发布CUA-HandCrafted基准,包含793个回合,覆盖24个复杂网页任务、56种攻击模板、8类攻击方式和4种系统提示配置。针对Claude Sonnet 4.6和GPT-5.4,多步攻击成功率为0/140(Clopper-Pearson 95%置信上限2.60%);提示消融实验表明该防御能力源于模型权重。然而,这种防护不具备泛化性:在同源代码智能体基准SkillBench上,相同模型可被手工技能注入攻破,最高成功率达100%。我们认为文献中高ASR主要归因于强化学习优化的注入文本,而非攻击类别本身。前沿安全防护具有领域特异性,尤其针对浏览器界面。若不公开优化过的攻击字符串,或将浏览器安全结论外推至其他模态,将导致发表的ASR数据无法复现。
原文摘要 · Abstract (English)
Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster on retired models and on the most-vulnerable model in each paper's panel. We ask whether those techniques, reproduced as hand-crafted templates, still work against current frontier CUAs. We release CUA-HandCrafted, a public benchmark of 793 episodes spanning 24 multi-step web tasks, 56 attack templates, 8 attack families, and 4 system-prompt configurations. Against Claude Sonnet 4.6 and GPT-5.4 we measure 0/140 multi-step attack success (Clopper-Pearson 95% upper bound 2.60%); a prompt ablation shows this resistance lives in the model weights. Yet it does not generalize: on a sister coding-agent benchmark (SkillBench), the same weights fall to hand-crafted skill-injection at up to 100%. We argue that the literature's high ASR is largely attributable to RL-optimized injection text rather than the attack categories, and that frontier safety hardening is domain-conditioned, specific to the heavily-targeted browser surface. Reporting techniques without releasing the optimized strings, or extrapolating browser-domain safety to other CUA modalities, makes published ASR numbers unreproducible.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。