教网页智能体识别欺骗性界面,提升安全性和可靠性。
Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces

- 用双阶段框架结合奖励学习与不对称惩罚,识别欺骗行为。
- 在1407个场景中使欺骗攻击成功率降低53.8%。
- 适用于需要高可靠性的自动化网页交互系统。
基于视觉-语言模型的网页智能体虽能自主操作图形界面,但仍易受欺骗性元素干扰。现有方法或仅检测欺骗而未融入任务,或仅记录攻击而不提供防御方案。本文首次形式化欺骗感知的网页智能体防御问题,提出DUDE(欺骗界面检测与评估)框架,采用混合奖励学习与非对称惩罚机制,并通过经验总结将失败模式提炼为可迁移的指导策略。构建了涵盖四个领域和多种欺骗类型的基准数据集RUC(Real UI Clickboxes),共1,407个场景。实验表明,DUDE在保持任务性能的同时,将欺骗敏感度降低53.8%,为部署鲁棒网页智能体奠定基础。
原文摘要 · Abstract (English)
Vision-language model (VLM) based web agents demonstrate impressive autonomous GUI interaction but remain vulnerable to deceptive interface elements. Existing approaches either detect deception without task integration or document attacks without proposing defenses. We formalize deception-aware web agent defense and propose DUDE (Deceptive UI Detector & Evaluator), a two-stage framework combining hybrid-reward learning with asymmetric penalties and experience summarization to distill failure patterns into transferable guidance. We introduce RUC (Real UI Clickboxes), a benchmark of 1,407 scenarios spanning four domains and deception categories. Experiments show DUDE reduces deception susceptibility by 53.8% while maintaining task performance, establishing an effective foundation for robust web agent deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。