构建跨市场多语言对话式商品搜索数据集,推动真实场景下的个性化推荐研究。
PSCon: Product Search Through Conversations
- 采用人类协作对话方式收集真实语料,保证语言自然性与多样性。
- 覆盖双市场、双语言,支持跨市场与多语言的对话搜索研究。
- 提供六项子任务基准,助力对话式搜索系统全面评估与优化。
对话式商品搜索(CPS)系统通过自然语言与用户交互,提供个性化和上下文感知的商品列表。然而,现有研究大多局限于模拟对话,因缺乏真实人类语言驱动的CPS数据集。此外,现有的电商对话数据集通常针对特定市场或语言,难以支持跨市场和多语言应用。本文提出一种CPS数据采集协议,构建了一个名为PSCon的新数据集,支持通过类人语言进行商品搜索。该数据集采用受训的人类-人类对话采集方式,覆盖双市场和两种语言。通过定义CPS任务,该数据集支持对六项子任务的深入研究:用户意图识别、关键词提取、系统动作预测、问题选择、商品排序和回复生成。同时,本文对数据集进行了详细分析,并提出了一个基准模型。所提出的数据集和模型将有助于推动未来对话式商品搜索的研究。
原文摘要 · Abstract (English)
Conversational Product Search ( CPS ) systems interact with users via natural language to offer personalized and context-aware product lists. However, most existing research on CPS is limited to simulated conversations, due to the lack of a real CPS dataset driven by human-like language. Moreover, existing conversational datasets for e-commerce are constructed for a particular market or a particular language and thus can not support cross-market and multi-lingual usage. In this paper, we propose a CPS data collection protocol and create a new CPS dataset, called PSCon, which assists product search through conversations with human-like language. The dataset is collected by a coached human-human data collection protocol and is available for dual markets and two languages. By formulating the task of CPS, the dataset allows for comprehensive and in-depth research on six subtasks: user intent detection, keyword extraction, system action prediction, question selection, item ranking, and response generation. Moreover, we present a concise analysis of the dataset and propose a benchmark model on the proposed CPS dataset. Our proposed dataset and model will be helpful for facilitating future research on CPS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。