零样本视觉搜索提升二手平台商品发现效率
Zero-Shot Retrieval for Scalable Visual Search in a Two-Sided Marketplace
- 采用多语言SigLIP模型实现无需训练的图像检索
- 相比基线,nDCG@5提升13.3%,交易率最高增40.9%
- 适合希望快速部署视觉搜索的电商平台团队
视觉搜索为用户探索多样商品目录提供直观方式,尤其适用于消费者对消费者(C2C)市场中信息不结构化、以视觉驱动的场景。本文介绍在Mercari C2C平台上部署的可扩展视觉搜索系统,该平台用户兼具买家与卖家身份。评估了近期视觉-语言模型在零样本图像检索中的表现,并与现有微调基线进行对比。系统整合实时推理与后台索引流程,通过降维优化统一嵌入管道。离线评估基于用户交互日志显示,多语言SigLIP模型在多个检索指标上均优于其他模型,nDCG@5相较基线提升13.3%。一周线上A/B测试进一步验证实际效果,实验组在参与度和转化率上显著提升,通过图像搜索的交易率最高增长40.9%。研究结果表明,近期零样本模型可作为生产环境的强基准,使团队以极低投入快速部署有效视觉搜索系统,同时保留未来基于数据或领域需求进行微调的灵活性。
原文摘要 · Abstract (English)
Visual search offers an intuitive way for customers to explore diverse product catalogs, particularly in consumer-to-consumer (C2C) marketplaces where listings are often unstructured and visually driven. This paper presents a scalable visual search system deployed in Mercari's C2C marketplace, where end-users act as buyers and sellers. We evaluate recent vision-language models for zero-shot image retrieval and compare their performance with an existing fine-tuned baseline. The system integrates real-time inference and background indexing workflows, supported by a unified embedding pipeline optimized through dimensionality reduction. Offline evaluation using user interaction logs shows that the multilingual SigLIP model outperforms other models across multiple retrieval metrics, achieving a 13.3% increase in nDCG@5 over the baseline. A one-week online A/B test in production further confirms real-world impact, with the treatment group showing substantial gains in engagement and conversion, up to a 40.9% increase in transaction rate via image search. Our findings highlight that recent zero-shot models can serve as a strong and practical baseline for production use, which enables teams to deploy effective visual search systems with minimal overhead, while retaining the flexibility to fine-tune based on future data or domain-specific needs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。