arXiv:2606.12608cs.CLcs.LG2026-06被引 1

构建首个专家设计的购物对话评估基准,测试模型多轮推理与决策能力。

Shopping Reasoning Bench: An Expert-Authored Benchmark for Multi-Turn Conversational Shopping Assistants

论文配图:Shopping Reasoning Bench: An Expert-Authored Benchmark for Multi-Turn Conversational Shopping Assistants
图 1 · 摘自论文原文
  • 基于零售专家撰写的525个购物任务,涵盖单轮与多轮对话
  • 九个模型平均通过率仅57%–77%,多轮任务表现显著下降
  • 揭示当前模型缺乏专家级购物建议能力,适合研究对话系统优化者

对话式购物助手已服务数亿用户,但现有基准未能联合评估开放性多轮推理、领域专长和标准级别质量。购物推理具有独特性:不同于事实问答或可验证代码生成,它需在主观偏好、预算约束和跨商品权衡间平衡,而现有电商与通用基准均缺乏此能力。我们推出购物推理基准(Shopping Reasoning Bench),包含525个任务(232个单轮,293个多轮),由零售专家撰写10863条重要性加权二元评判标准。这些标准按五类推理与十五个子类组织,覆盖偏好细化、权衡分析和兼容性评估等需求。对九个模型(三类:GPT、Claude、Gemini)的评估显示,整体通过率仅为57%–77%。多轮任务中,所有模型在可选“超越要求”标准上得分比必选标准低13–29分,且随着对话推进性能下降4–18分。结果表明当前模型仅能完成基础购物协助,难以达到专家级建议水平,使该基准成为未来购物助手研发的重要挑战平台。

原文摘要 · Abstract (English)

Conversational shopping assistants now serve hundreds of millions of customers, yet no existing benchmark jointly evaluates the open-ended multi-turn reasoning, domain expertise, and criterion-level quality that real shopping conversations demand. Shopping reasoning is unique among language model applications. Unlike factual question answering or verifiable code generation, it requires balancing subjective preferences, budget constraints, and cross-product trade-offs across multi-turn dialogue, capabilities absent from previous e-commerce and general-purpose benchmarks. We introduce the Shopping Reasoning Bench, an expert-authored benchmark of 525 missions (232 single-turn, 293 multi-turn) with 10863 importance-weighted binary rubrics authored by retail domain experts. These criteria are organized under a taxonomy of five reasoning categories and fifteen subcategories covering diverse demands such as preference refinement, trade-off analysis, and compatibility assessment. An evaluation of nine models across three families (GPT, Claude, Gemini) shows that pass rates reach only 57--77% overall. On multi-turn missions, all models score 13--29 points lower on optional above-and-beyond criteria than on required ones, and performance degrades 4--18 points as conversations progress. These gaps show that current models handle basic shopping assistance but fall short of expert-level advice, making Shopping Reasoning Bench a challenging testbed for future shopping assistant development.

对话系统购物助手多轮推理评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。