新基准Amazon-Bench评估电商智能体功能与安全,覆盖搜索之外的复杂操作。
A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains
- 基于网页内容生成覆盖多种功能的用户查询,如地址管理、心愿单操作。
- 首次同时评估智能体任务完成率与潜在风险,发现现有模型存在安全漏洞。
- 适合研究电商智能体可靠性与安全性的人群使用。
网络智能体在电商网站上展现出巨大潜力,但现有基准存在两大缺陷:一是仅关注产品搜索任务(如查找Apple Watch),忽略账户管理、礼品卡等真实平台功能;二是仅评估任务完成度,忽视潜在风险。实践中,智能体可能误购商品、删除保存地址或错误配置自动续费。为解决此问题,我们提出新基准Amazon-Bench。通过分析网页内容与交互元素(如按钮、复选框)构建数据生成管道,生成涵盖地址管理、心愿单操作、品牌店铺关注等多样化的功能导向查询。同时设计自动化评估框架,同步衡量智能体性能与安全性。系统评估显示,当前智能体在复杂任务中表现不佳且存在安全隐患,凸显开发更鲁棒、可靠智能体的必要性。
原文摘要 · Abstract (English)
Web agents have shown great promise in performing many tasks on ecommerce website. To assess their capabilities, several benchmarks have been introduced. However, current benchmarks in the e-commerce domain face two major problems. First, they primarily focus on product search tasks (e.g., Find an Apple Watch), failing to capture the broader range of functionalities offered by real-world e-commerce platforms such as Amazon, including account management and gift card operations. Second, existing benchmarks typically evaluate whether the agent completes the user query, but ignore the potential risks involved. In practice, web agents can make unintended changes that negatively impact the user account or status. For instance, an agent might purchase the wrong item, delete a saved address, or incorrectly configure an auto-reload setting. To address these gaps, we propose a new benchmark called Amazon-Bench. To generate user queries that cover a broad range of tasks, we propose a data generation pipeline that leverages webpage content and interactive elements (e.g., buttons, check boxes) to create diverse, functionality-grounded user queries covering tasks such as address management, wish list management, and brand store following. To improve the agent evaluation, we propose an automated evaluation framework that assesses both the performance and the safety of web agents. We systematically evaluate different agents, finding that current agents struggle with complex queries and pose safety risks. These results highlight the need for developing more robust and reliable web agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。