arXiv:2506.02839cs.IRcs.AI2025-06被引 41

构建真实购物场景的评测基准,提升网购代理评估真实性

DeepShop: A Benchmark for Deep Research Shopping Agents

  • 从真实用户查询出发,生成多维度复杂搜索需求
  • 分三等级评估代理在筛选与排序上的表现,最高成功率仅41%
  • 适合研究智能导购、网页代理与复杂任务理解的团队

面向在线购物的网络代理在自动化电商交互方面展现出巨大潜力。现有评测基准难以反映真实购物场景的复杂性,多采用简单查询和确定性路径,如“找iPhone 15”。真实购物涉及多维商品属性、搜索筛选条件及用户个性化排序偏好。为此,我们提出DeepShop,一个用于评估复杂真实网购环境下的网络代理的基准。DeepShop包含三个核心部分:(1) 查询多样性演化:基于真实用户查询,生成五个热门电商领域的多样化查询;(2) 查询复杂度演化:综合商品属性、筛选项和排序偏好,将查询分为易、中、难三级;(3) 细粒度与整体评估:设计自动化评估框架,分别评估属性、筛选、排序等细粒度表现,并报告整体成功率。对检索增强生成(RAG)、网络代理与深度研究系统进行系统评估,结果显示RAG因缺乏网页交互能力,在复杂查询上表现不佳;其他方法在筛选与排序上仍面临显著挑战,整体成功率最低仅41%。通过跨品类、复杂度分析与错误归因,为深化购物代理研究提供支持。

原文摘要 · Abstract (English)

Web agents for online shopping have shown great promise in automating user interactions across e-commerce platforms. Benchmarks for assessing such agents do not reflect the complexity of real-world shopping scenarios, as they often consist of overly simple queries with deterministic paths, such as "Find iPhone 15." Real shopping scenarios are inherently more layered, involving multi-dimensional product attributes, search filters, and user-specific sorting preferences. To address this gap, we introduce DeepShop, a benchmark designed to evaluate web agents in complex and realistic online shopping environments. DeepShop comprises three key components. (1) Query diversity evolution: Starting from real user queries, we generate diverse queries across five popular online shopping domains. (2) Query complexity evolution: We further evolve these queries to increase complexity, considering product attributes, search filters, and sorting preferences, and classify them into three levels: easy, medium, and hard, based on the number of evolutions. (3) Fine-grained and holistic evaluation: We propose an automated evaluation framework that assesses agent performance in terms of fine-grained aspects (product attributes, search filters, and sorting preferences) and reports the overall success rate through holistic evaluation. We conduct a systematic evaluation of retrieval-augmented generation (RAG) methods, web agents, and deep research systems. Results show that RAG struggles with complex queries due to its lack of web interaction, while other methods face significant challenges with filters and sorting preferences, leading to low overall success rates. We also perform cross-category, complexity-based evaluations and error analyses to support the advancement of deep research shopping agents.

购物代理评测基准复杂查询自动化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。