arXiv:2608.09282cs.AIcs.CL2026-08

评测大模型在有预算和优惠券限制下的组合购物能力。

ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

论文配图:ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons
图 1 · 摘自论文原文
  • 构建模拟电商环境,生成可验证的组合购物任务
  • 多轮验证确保订单合规性与优惠最优
  • 适合研究大模型在复杂消费场景中的推理能力

现实购物常需组合购买互补商品,如设备配置、餐食准备、活动策划和多人外卖。这类任务需联合推理商品兼容性、库存、门店要求、配送费、优惠券和预算。现有评估方法难以处理多个合理答案的情况,且无法检测无效订单或错误支付。我们提出ComboShoppingBench,一个面向开放但可验证的组合购物基准,通过探索代理生成合法且语义一致的商品篮子,据此合成优惠券、预算约束、用户查询及对齐评估标准。评估阶段由大模型判断语义满足度与响应质量,同时通过确定性验证检查商品ID有效性、预算合规性和优惠券最优性。实验表明,即使强模型在该基准上仍表现不佳,凸显了可靠、约束感知组合购物的巨大改进空间。

原文摘要 · Abstract (English)

Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise in device setup, meal preparation, event planning, and group takeout ordering, requiring joint reasoning about item compatibility, availability, store-level requirements, delivery fees, coupons, and budgets. Evaluation is challenging because multiple baskets may satisfy the same request, making exact-match metrics unsuitable, whereas semantic evaluation alone cannot detect infeasible orders, invalid coupon combinations, or incorrect payments. We introduce ComboShoppingBench, an agentic shopping benchmark for open-ended yet verifiable basket construction in a simulated commerce and takeout environment. During task synthesis, an exploration agent constructs a feasible and semantically coherent basket of purchasable products; this witness guides the generation of coupons, budget constraints, user queries, and aligned evaluation rubrics. During evaluation, LLM judges assess semantic satisfaction, response quality, and claim faithfulness, while deterministic validation checks product-ID validity, budget compliance, and coupon optimality. Experiments with diverse LLM agents demonstrate that even strong agents struggle on ComboShoppingBench, highlighting substantial room for improvement in reliable, constraint-aware combo shopping.

大模型评估组合购物约束推理智能导购

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。