arXiv:2511.11305cs.IRcs.AI2025-11被引 5

MOON提升电商搜索广告点击率,通过多模态表示学习实现20%增长。

MOON Embedding: Multimodal Representation Learning for E-commerce Search Advertising

  • 分三阶段训练:预训练、后训练、应用,融合多模态信息
  • 在线点击率提升20.00%,三年五次迭代持续优化
  • 揭示图像召回是关键中间指标,适合推荐系统研究者参考

我们提出MOON,一套面向电商应用的可持续迭代式多模态表示学习体系。该系统已全面部署于淘宝搜索广告系统,涵盖检索、相关性、排序等环节。在点击率(CTR)预测任务上实现整体+20.00%的线上提升。过去三年中,该项目成为CTR预测任务上最大改进,并完成五次全规模迭代。通过探索与迭代,我们积累了宝贵经验。MOON采用“预训练、后训练、应用”三阶段范式,有效整合多模态表示与下游任务。为弥合多模态学习目标与下游训练之间的错位,我们引入“汇率”概念,量化中间指标改进对下游效果的转化效率。分析发现,基于图像的搜索召回是关键中间指标,指导模型优化。三年间,MOON在数据处理、训练策略、模型架构和下游应用四个维度持续演进。此外,我们还系统研究了电商场景下的缩放规律,考察训练样本量、负样本数量及用户行为序列长度等因素的影响。

原文摘要 · Abstract (English)

We introduce MOON, our comprehensive set of sustainable iterative practices for multimodal representation learning for e-commerce applications. MOON has already been fully deployed across all stages of Taobao search advertising system, including retrieval, relevance, ranking, and so on. The performance gains are particularly significant on click-through rate (CTR) prediction task, which achieves an overall +20.00% online CTR improvement. Over the past three years, this project has delivered the largest improvement on CTR prediction task and undergone five full-scale iterations. Throughout the exploration and iteration of our MOON, we have accumulated valuable insights and practical experience that we believe will benefit the research community. MOON contains a three-stage training paradigm of "Pretraining, Post-training, and Application", allowing effective integration of multimodal representations with downstream tasks. Notably, to bridge the misalignment between the objectives of multimodal representation learning and downstream training, we define the exchange rate to quantify how effectively improvements in an intermediate metric can translate into downstream gains. Through this analysis, we identify the image-based search recall as a critical intermediate metric guiding the optimization of multimodal models. Over three years and five iterations, MOON has evolved along four critical dimensions: data processing, training strategy, model architecture, and downstream application. The lessons and insights gained through the iterative improvements will also be shared. As part of our exploration into scaling effects in the e-commerce field, we further conduct a systematic study of the scaling laws governing multimodal representation learning, examining multiple factors such as the number of training tokens, negative samples, and the length of user behavior sequences.

多模态推荐系统电商表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。