arXiv:2605.16116cs.AI2026-05被引 3

构建可复现的电商网页智能体评测框架,兼顾真实性和可控性。

ShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E-Commerce Web Agents

论文配图:ShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E-Commerce Web Agents
图 1 · 摘自论文原文
  • 用匿名化数据生成独立沙盒店铺,保持真实结构
  • 224个任务验证:合成店与真实店表现相关性高
  • 适合研究电商智能体评估的学者和开发者

开发与评估电商网页智能体需要既能保留有意义任务结构,又具备可控、可复现、可扩展特性的环境。现有方法存在权衡:真实店铺虽真实但不稳定、难观测、不可复现;人工搭建的沙盒环境虽可控,却仅覆盖有限的页面布局、商品目录、策略和交互模式。我们指出核心瓶颈在于方法论——缺乏同时满足真实、多样、可控、可检视、可复现的规模化构建方式。为此提出 ShopGym 框架,包含两层:ShopArena 将真实店铺转化为自包含沙盒,通过匿名化规格与分阶段验证生成过程;ShopGuru 基于商品目录、导航结构、策略与交互能力,在七类技能上合成基准任务。二者结合生成稳定、可重置、可检视的评测样本,保留购物任务的关键结构与评估信号。通过图结构分析及基于智能体的行为评估,我们在六个沙盒店铺(三组合成数据、三组真实数据)中验证了224项任务,结果表明合成店铺保留了真实店铺的关键结构特征,且智能体在合成店上的表现与真实店正相关。

原文摘要 · Abstract (English)

Developing and evaluating e-commerce web agents requires environments that preserve meaningful task structure while enabling controllable, reproducible, and scalable scientific comparison. Existing methodologies force a tradeoff: live storefronts provide realism but are non-stationary, difficult to inspect, and irreproducible, while hand-built sandbox benchmarks provide control but cover only a narrow range of layouts, catalogs, policies, and interaction patterns. We argue that the core bottleneck is methodological: the field lacks a scalable way to construct evaluation settings that are simultaneously realistic, diverse, controllable, inspectable, and reproducible. We introduce ShopGym, an integrated framework for realistic simulation and scalable benchmarking of e-commerce web agents. ShopGym is a framework for constructing e-commerce simulation environments and grounded benchmark tasks. Its simulation layer, ShopArena, converts live seed storefronts into self-contained sandbox shops through anonymized shop specifications and a staged, validated generation process. On top of these simulated storefronts, ShopGuru synthesizes benchmark tasks across seven skill categories, grounding each task in the shop's catalog, navigation structure, policies, and interaction affordances. Together, ShopArena and ShopGuru produce self-contained, resettable, inspectable, and stable evaluation artifacts that preserve structural properties and agent-evaluation signals relevant to shopping tasks. We validate the framework through graph-based structural analysis and agent-based behavioral evaluation with 224 generated tasks across six sandbox shops: three constructed with synthetic data and three with real data. Our results show that the synthetic shops preserve key structural properties of live storefronts, with agent performance on synthetic shops positively correlated with performance on live storefronts.

电商智能体仿真评测可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。