arXiv:2504.07110cs.IRcs.LG2025-04被引 2

用多模态模型生成门徒商品与用户意图的语义嵌入,提升推荐效果。

DashCLIP: Leveraging multimodal models for generating semantic embeddings for DoorDash

  • 通过对比学习对齐单模态与多模态编码器,联合训练商品和查询编码器。
  • 在商品分类与相关性预测任务中表现优异,点击率与转化率显著提升。
  • 无需依赖用户行为历史,适合电商场景的通用嵌入构建。

尽管视觉语言模型在各类生成任务中取得成功,但获取高质量的产品与用户意图语义表示仍具挑战,因现成模型难以捕捉实体间的细微关系。本文提出一种联合训练框架,通过图像-文本数据上的对比学习对齐单模态与多模态编码器,训练一个由大语言模型标注的相关性数据集驱动的查询编码器,摆脱对用户交互历史的依赖。这些嵌入表现出强泛化能力,在商品分类、相关性预测等应用中表现良好。在个性化广告推荐中,部署后点击率与转化率显著提升,验证了其对关键业务指标的影响。我们认为该框架的灵活性为丰富电商业务用户体验提供了有力解决方案。

原文摘要 · Abstract (English)

Despite the success of vision-language models in various generative tasks, obtaining high-quality semantic representations for products and user intents is still challenging due to the inability of off-the-shelf models to capture nuanced relationships between the entities. In this paper, we introduce a joint training framework for product and user queries by aligning uni-modal and multi-modal encoders through contrastive learning on image-text data. Our novel approach trains a query encoder with an LLM-curated relevance dataset, eliminating the reliance on engagement history. These embeddings demonstrate strong generalization capabilities and improve performance across applications, including product categorization and relevance prediction. For personalized ads recommendation, a significant uplift in the click-through rate and conversion rate after the deployment further confirms the impact on key business metrics. We believe that the flexibility of our framework makes it a promising solution toward enriching the user experience across the e-commerce landscape.

多模态嵌入推荐系统门徒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。