首个离线多店铺购物基准,评估大模型网购代理的复杂比价能力。
WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents
- 构建四个模拟店铺,数据来自Common Crawl,支持多源异构商品
- 任务涵盖比价、互补品搜索等,最佳代理完成率不足65%
- 适合评估复杂网络代理在真实电商场景下的泛化能力
基于大语言模型的网络代理有望自动化完成跨多电商平台的长周期任务,如寻找满足需求的最便宜商品。现有评测基准或需在线执行(如DeepShop、ShoppingComp),或仅覆盖单店简单任务(如WebShop、WebArena、Mind2Web)。缺乏能模拟多店铺、异构商品数据并要求复杂检索的离线评测平台。为此,我们提出WebMall,首个离线多店铺基准,用于评估网络代理在复杂比价任务中的表现。WebMall包含四个由Common Crawl提取商品数据构建的模拟店铺,任务范围从精准搜索、价格比较,到互补品或替代品的高级搜索及结账流程。使用八种不同配置的代理进行验证,结果显示最佳代理在最便宜商品搜索和模糊搜索任务中完成率均低于65%,凸显该基准的挑战性。
原文摘要 · Abstract (English)
LLM-based web agents have the potential to automate long-running web tasks, such as searching for products in multiple e-shops and subsequently ordering the cheapest products that meet the users needs. Benchmarks for evaluating web agents either require agents to perform tasks online using the live Web or offline using simulated environments, the latter allowing for the exact reproduction of the experimental setup. While DeepShop and ShoppingComp provide online benchmarks that require agents to perform challenging shopping tasks, existing offline benchmarks such as WebShop, WebArena, and Mind2Web cover only comparatively simple e-commerce tasks performed against a single shop containing product data from a single source. What is missing is an e-commerce benchmark that simulates multiple shops containing heterogeneous product data and requires agents to perform complex retrieval tasks. We fill this gap by introducing WebMall, the first offline multi-shop benchmark for evaluating web agents on challenging comparison shopping tasks. WebMall consists of four simulated shops populated with product data extracted from the Common Crawl. The WebMall tasks range from specific product searches and price comparisons to advanced searches for complementary or substitute products, as well as checkout processes. We validate WebMall using eight agents that differ in observation space, availability of short-term memory, and the employed LLM. The validation highlights the difficulty of the benchmark, with the best-performing agents achieving task completion rates below 65% in the task categories cheapest product search and vague product search.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。