arXiv:2510.21671cs.IR2025-10

通过数据重构提升多语言电商搜索效果,无需改模型

A Data-Centric Approach to Multilingual E-Commerce Product Search: Case Study on Query-Category and Query-Item Relevance

  • 用翻译扩充、语义负样本、自验证过滤三策略优化训练数据
  • 在CIKM AnalytiCup 2025上显著提升F1分数,超越强基线
  • 适合实际电商场景中资源有限的多语言搜索系统构建

多语言电商搜索面临语言间数据不平衡、标签噪声及低资源语言监督不足等问题,制约了相关性模型的跨语言泛化能力。本文提出一种与架构无关的数据中心框架,聚焦两个核心任务:查询-类别(QC)相关性和查询-商品(QI)相关性。不修改模型,而是通过三种互补策略重构训练数据:(1) 基于翻译的增强,为训练中缺失的语言合成样本;(2) 语义负采样,生成难负例并缓解类别不平衡;(3) 自验证过滤,检测并剔除可能误标的数据。在CIKM AnalytiCup 2025数据集上评估,该方法持续显著提升F1分数,优于多个强基线模型,并在官方竞赛中取得竞争力结果。研究表明,系统性的数据工程可与复杂模型改进媲美,且更易部署,为真实电商场景中的鲁棒多语言搜索系统提供可操作指导。

原文摘要 · Abstract (English)

Multilingual e-commerce search suffers from severe data imbalance across languages, label noise, and limited supervision for low-resource languages--challenges that impede the cross-lingual generalization of relevance models despite the strong capabilities of large language models (LLMs). In this work, we present a practical, architecture-agnostic, data-centric framework to enhance performance on two core tasks: Query-Category (QC) relevance (matching queries to product categories) and Query-Item (QI) relevance (matching queries to product titles). Rather than altering the model, we redesign the training data through three complementary strategies: (1) translation-based augmentation to synthesize examples for languages absent in training, (2) semantic negative sampling to generate hard negatives and mitigate class imbalance, and (3) self-validation filtering to detect and remove likely mislabeled instances. Evaluated on the CIKM AnalytiCup 2025 dataset, our approach consistently yields substantial F1 score improvements over strong LLM baselines, achieving competitive results in the official competition. Our findings demonstrate that systematic data engineering can be as impactful as--and often more deployable than--complex model modifications, offering actionable guidance for building robust multilingual search systems in the real-world e-commerce settings.

多语言搜索数据增强电商推荐标签噪声

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。