构建真实购物场景的智能体评估基准,挑战复杂用户意图理解。
ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
- 基于真实商品构建多层级购物意图模拟框架。
- 大语言模型在任务中成功率不足50%。
- 适合研究电商智能体与小模型蒸馏的学者。
现有电商平台评测主要聚焦于基础用户意图(如查找或购买商品),但现实用户常面临更复杂的任务,如使用优惠券、管理预算、寻找多商品卖家等。为填补这一差距,我们提出 ShoppingBench——一个端到端的真实世界购物基准,涵盖日益复杂的具身意图。我们设计了一种可扩展框架,基于真实商品采样生成多样化用户指令。为确保评估一致性与可靠性,我们构建了一个包含超过250万真实商品的大规模购物沙箱环境,支持交互式仿真。实验表明,即使最先进的语言智能体(如 GPT-4.1)在本基准任务中的绝对成功率也低于50%,凸显了该基准的挑战性。此外,我们提出轨迹蒸馏策略,结合监督微调与合成轨迹上的强化学习,将大语言模型的能力迁移到小型模型。最终训练出的代理在性能上可媲美 GPT-4.1。
原文摘要 · Abstract (English)
Existing benchmarks in e-commerce primarily focus on basic user intents, such as finding or purchasing products. However, real-world users often pursue more complex goals, such as applying vouchers, managing budgets, and finding multi-products seller. To bridge this gap, we propose ShoppingBench, a novel end-to-end shopping benchmark designed to encompass increasingly challenging levels of grounded intent. Specifically, we propose a scalable framework to simulate user instructions based on various intents derived from sampled real-world products. To facilitate consistent and reliable evaluations, we provide a large-scale shopping sandbox that serves as an interactive simulated environment, incorporating over 2.5 million real-world products. Experimental results demonstrate that even state-of-the-art language agents (such as GPT-4.1) achieve absolute success rates under 50% on our benchmark tasks, highlighting the significant challenges posed by our ShoppingBench. In addition, we propose a trajectory distillation strategy and leverage supervised fine-tuning, along with reinforcement learning on synthetic trajectories, to distill the capabilities of a large language agent into a smaller one. As a result, our trained agent achieves competitive performance compared to GPT-4.1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。