构建首个长周期电商任务基准,提升大模型购物代理的偏好理解能力。
Shopping Companion: Benchmarking and Training LLM Agents for Long-Horizon Preference-Grounded E-Commerce Tasks
- 设计基于真实商品库的跨会话偏好记忆任务,模拟复杂购物场景。
- 40亿参数轻量模型在任务中表现优于强基线,成功率超70%。
- 无需标注的工具奖励机制,解决长程任务中的奖励稀疏问题。
在电子商务中,大语言模型代理在推荐、预算管理与组合优惠等任务中展现出潜力,准确捕捉长期对话中的用户偏好至关重要。然而,进展受限于两大挑战:(1) 缺乏评估长周期偏好感知购物任务的基准;(2) 缺少细粒度的监督信号用于代理训练。为填补基准空白,我们提出 Shopping Companion Bench,一个包含两个需跨会话偏好记忆任务的新基准,基于超过120万条真实商品数据。分析发现,主要失败原因包括由偏好幻觉引发的级联错误,以及对产品属性与用户需求验证不足。为此,我们设计无标注、按工具分项的奖励机制,为每一步工具调用提供过程监督,缓解长程任务中的奖励稀疏问题。实验表明,即使最先进的GPT-5模型在该基准上的成功率为70%以下,凸显任务难度。值得注意的是,我们微调的40亿参数轻量模型在偏好捕获与任务表现上持续优于强基线,验证了奖励设计的有效性。
原文摘要 · Abstract (English)
In e-commerce, LLM agents show promise for shopping tasks such as recommendations, budget management, and bundle deals, where accurately capturing user preferences from long-horizon conversations is critical. However, progress is limited by two key challenges: (1) the absence of benchmarks for evaluating long-term preference-aware shopping tasks, and (2) the lack of fine-grained supervision for shopping agent training. To fill the benchmark gap, we introduce Shopping Companion Bench, a novel benchmark comprising two shopping tasks that require cross-session preference memory, grounded in a product pool of over 1.2 million real-world items. Our analysis further identifies two major sources of failure on this benchmark: cascading errors caused by preference hallucination, and insufficient verification of product attributes against user requirements. To address these failure modes, we design annotation-free, tool-wise rewards that provide process supervision for each tool call, alleviating reward sparsity in long-horizon tasks. Experimental results demonstrate that even state-of-the-art models such as GPT-5 achieve success rates below 70%, highlighting the difficulty of our benchmark. Notably, our fine-tuned lightweight 4B model consistently outperforms strong baselines in both preference capture and task performance, suggesting the effectiveness of our reward design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。