首个基于真实购物日志的多模态购物助手评测基准,支持图文多轮交互。
MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents

- 基于真实用户日志构建,含图文多轮对话与购买意图标注。
- 模型需从图像和对话中联合推断需求,匹配满足所有条件的商品。
- 提供可复现的离线沙箱和微调数据集,助力开源模型性能提升。
在线消费者越来越多地依赖AI购物助手,通过图片和多轮对话表达难以用文字描述的产品需求。然而,现有基准大多依赖纯文本或合成请求,无法充分反映图文结合的真实购物场景。我们提出MMShopBench,首个基于真实购物日志的多模态、多轮购物代理评测基准。该数据集由精心清洗并人工标注的购物日志构建,包含每条请求的购买意图和强制性产品要求。模型需联合解析用户图像与多轮对话以推断需求,通过图文搜索召回候选商品,并利用商品图像与结构化属性验证其是否满足全部要求。我们采用基于证据的多模态评估协议,评测主流开源与专有模型,并构建配套训练集用于微调开源模型。为保障实验可复现性,我们搭建了离线购物沙箱,微调后开源模型性能显著接近领先专有模型,证明了训练数据的有效性。
原文摘要 · Abstract (English)
Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in text alone. However, existing benchmarks largely rely on text-only or synthetic requests, underrepresenting complex real-world shopping requirements jointly expressed through images and language. We introduce MMShopBench, the first real-log benchmark for multimodal, multi-turn shopping agents. Built from carefully cleaned and manually annotated shopping logs, MMShopBench provides ground-truth annotations of each request's purchase intent and mandatory product requirements. Agents must infer these requirements jointly from user images and multi-turn dialogue, retrieve candidate products through image and text search, and verify that each candidate satisfies all requirements using its product images and structured attributes. We evaluate representative open-source and proprietary models using an evidence-grounded multimodal protocol and construct a companion training set for fine-tuning an open-source model. To ensure reproducible experimentation, we build an offline shopping sandbox, where fine-tuning substantially narrows the performance gap between our open-source model and leading proprietary models, demonstrating the effectiveness of our training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。