arXiv:2510.20674cs.IRcs.CL2025-10

团队通过多语言数据增强与模型微调,提升电商搜索匹配精度。

Analyticup E-commerce Product Search Competition Technical Report from Team Tredence_AICOE

  • 用翻译扩充数据,覆盖所有目标语言,实现跨语种训练。
  • Gemma-3 12B(4-bit)在两项任务中分别取得最佳效果,平均F1达0.8857。
  • 适合关注多语言电商搜索与大模型轻量化应用的研究者。

本文介绍由Tredence_AICOE团队开发的多语言电商搜索系统。竞赛包含两项多语言相关性任务:查询-类别(QC)相关性,评估用户查询与商品类别的匹配程度;查询-商品(QI)相关性,衡量多语言查询与具体商品列表的匹配度。为确保全语言覆盖,团队将现有数据集翻译至开发集中缺失的语言,从而实现所有目标语言的训练。采用多种策略对Gemma-3 12B和Qwen-2.5 14B模型进行微调。其中,使用原始数据与翻译数据的Gemma-3 12B(4-bit)模型在QC任务中表现最优;结合原始、翻译及少数类数据的版本在QI任务中表现最佳。该方法在最终排行榜上获得第4名,私有测试集平均F1得分为0.8857。

原文摘要 · Abstract (English)

This study presents the multilingual e-commerce search system developed by the Tredence_AICOE team. The competition features two multilingual relevance tasks: Query-Category (QC) Relevance, which evaluates how well a user's search query aligns with a product category, and Query-Item (QI) Relevance, which measures the match between a multilingual search query and an individual product listing. To ensure full language coverage, we performed data augmentation by translating existing datasets into languages missing from the development set, enabling training across all target languages. We fine-tuned Gemma-3 12B and Qwen-2.5 14B model for both tasks using multiple strategies. The Gemma-3 12B (4-bit) model achieved the best QC performance using original and translated data, and the best QI performance using original, translated, and minority class data creation. These approaches secured 4th place on the final leaderboard, with an average F1-score of 0.8857 on the private test set.

多语言搜索电商推荐大模型微调数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。