跨市场多语言商品问答新任务,用别国数据提升本国购物决策支持
Unlocking Markets: A Multilingual Benchmark to Cross-Market Question Answering

- 利用多语言资源丰富的市场数据,辅助回答主市场用户问题
- 覆盖17个市场、11种语言,超700万条商品相关问题
- 实验表明跨市场信息能显著提升问答和排序效果
用户在电商平台发布大量商品相关问题,影响购买决策。产品相关问答(PQA)旨在利用商品资源精准回应用户问题。本文提出多语言跨市场商品问答(MCPQA)新任务,定义为在主市场中通过另一个资源丰富的辅助市场,在多语言环境下提供商品问题的答案。我们构建了一个大规模数据集,涵盖17个市场、11种语言的超过700万条问题。对电子品类进行自动翻译,形成McMarket数据集。聚焦两个子任务:基于评论的答案生成与商品相关问题排序。使用大模型标注部分McMarket数据,并通过人工评估验证标注质量。在单市场与跨市场场景下,测试从传统词法模型到大模型的多种方法。结果表明,引入跨市场信息可显著提升两类任务性能。
原文摘要 · Abstract (English)
Users post numerous product-related questions on e-commerce platforms, affecting their purchase decisions. Product-related question answering (PQA) entails utilizing product-related resources to provide precise responses to users. We propose a novel task of Multilingual Cross-market Product-based Question Answering (MCPQA) and define the task as providing answers to product-related questions in a main marketplace by utilizing information from another resource-rich auxiliary marketplace in a multilingual context. We introduce a large-scale dataset comprising over 7 million questions from 17 marketplaces across 11 languages. We then perform automatic translation on the Electronics category of our dataset, naming it as McMarket. We focus on two subtasks: review-based answer generation and product-related question ranking. For each subtask, we label a subset of McMarket using an LLM and further evaluate the quality of the annotations via human assessment. We then conduct experiments to benchmark our dataset, using models ranging from traditional lexical models to LLMs in both single-market and cross-market scenarios across McMarket and the corresponding LLM subset. Results show that incorporating cross-market information significantly enhances performance in both tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。