OneVision用生成式框架统一多视角电商搜索,提升效率与转化。
OneVision: An End-to-End Generative Framework for Multi-view E-commerce Vision Search
- 构建视觉对齐的残差量化编码,统一多视角物体表示
- 实现离线性能持平线上,推理效率提升21%;在线点击率+2.15%
- 适合追求高效个性化搜索的电商平台使用
传统视觉搜索采用多阶段级联架构(MCA),在特征提取、召回、预排序和排序阶段间存在多视角表征差异与优化目标冲突,难以兼顾用户体验与转化率。本文提出端到端生成框架OneVision,基于视觉对齐的残差量化编码(VRQ),有效对齐同一物体在不同视角下的表征,同时保留产品独特特征。通过多阶段语义对齐机制,在保持强视觉相似性先验的同时融入用户偏好信息,实现个性化推荐。离线评估显示,OneVision性能媲美线上MCA,推理效率提升21%;A/B测试中,点击率提升2.15%,转化率提升2.27%,订单量增长3.12%。结果表明,以语义ID为中心的生成架构可统一检索与个性化,简化服务路径。
原文摘要 · Abstract (English)
Traditional vision search, similar to search and recommendation systems, follows the multi-stage cascading architecture (MCA) paradigm to balance efficiency and conversion. Specifically, the query image undergoes feature extraction, recall, pre-ranking, and ranking stages, ultimately presenting the user with semantically similar products that meet their preferences. This multi-view representation discrepancy of the same object in the query and the optimization objective collide across these stages, making it difficult to achieve Pareto optimality in both user experience and conversion. In this paper, an end-to-end generative framework, OneVision, is proposed to address these problems. OneVision builds on VRQ, a vision-aligned residual quantization encoding, which can align the vastly different representations of an object across multiple viewpoints while preserving the distinctive features of each product as much as possible. Then a multi-stage semantic alignment scheme is adopted to maintain strong visual similarity priors while effectively incorporating user-specific information for personalized preference generation. In offline evaluations, OneVision performs on par with online MCA, while improving inference efficiency by 21% through dynamic pruning. In A/B tests, it achieves significant online improvements: +2.15% item CTR, +2.27% CVR, and +3.12% order volume. These results demonstrate that a semantic ID centric, generative architecture can unify retrieval and personalization while simplifying the serving pathway.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。