让商品标题生成更符合用户搜索和购物意图的伪查询。
SAM-D2Q: Aligning Multimodal Doc2Query with Search Demand and Conversion for E-commerce

- 融合图文信息,用强化学习对齐电商搜索目标
- 在阿里速卖通上提升GMV 3.38%、支付数2.27%
- 适合电商搜索优化与多模态生成任务
电商平台搜索常因用户查询与商家商品标题间存在词汇不匹配问题,导致短标题无法覆盖用户多样表达或视觉属性。尽管文档转查询(Doc2Query)可生成伪查询扩展文档,但传统方法仅依赖文本且未针对电商业务目标优化,可能生成语义合理却商业无效的扩展,遗漏图像中的关键属性。为此,我们提出面向电商搜索的多模态文档扩展框架SAM-D2Q,其包含三个阶段:(1) 针对任务的多模态监督微调,增强商品标题、图片与用户查询的视觉语言理解;(2) 多模态数据增强,提升对关键视觉属性的感知与扩展覆盖;(3) 基于强化学习的偏好对齐,引导模型生成更契合用户意图与商业价值的伪查询。离线实验表明,SAM-D2Q显著优于传统Doc2Query方法。部署于阿里速卖通生产搜索系统后,线上业务指标提升,GMV增长+3.38%,支付数提升+2.27%。
原文摘要 · Abstract (English)
E-commerce search often suffers from vocabulary mismatch between user queries and merchant-authored product titles, since short titles cannot fully cover diverse user expressions or visual product attributes. Although Doc2Query alleviates this issue by generating pseudo-queries for document expansion, traditional methods are text-only and not optimized for e-commerce business objectives. As a result, they may produce semantically plausible but commercially ineffective expansions and miss key attributes present in product images. To this end, we propose E-commerce Search-Aligned Multimodal Doc2Query (SAM-D2Q), a business-aligned multimodal document expansion framework for e-commerce search under Boolean retrieval constraints. SAM-D2Q consists of three stages: (1) task-adapted multimodal supervised fine-tuning to enhance vision-language understanding of product titles, images, and user queries; (2) multimodal data augmentation to improve perception of key visual attributes and expansion coverage; and (3) reinforcement-learning-based preference alignment toward search business objectives, encouraging the model to generate pseudo-queries that better match user intent and commercial value. Offline experiments show that SAM-D2Q substantially improves retrieval performance over traditional Doc2Query methods. Deployed in the AliExpress production search system, SAM-D2Q improves online business metrics, increasing GMV by +3.38% and Pay Count by +2.27%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。