arXiv:2602.13704cs.IRcs.AI2026-02被引 3

将多模态检索精度与实时性结合,提升电商搜索体验

Pailitao-VL: Unified Embedding and Reranker for Real-Time Multi-Modal Industrial Search

  • 用全局语义原型替代对比学习,实现更精准的嵌入表示
  • 通过分块对比与校准评分,显著降低重排延迟
  • 在阿里电商平台验证,兼具高精度与高效率

本文提出Pailitao-VL,一个面向高精度、实时工业级多模态检索的综合性系统。针对现有SOTA方案在检索粒度不足、易受环境噪声干扰、效率与性能矛盾突出三大挑战,我们实现两大范式转变:首先,将嵌入范式从传统对比学习转向绝对ID识别任务,通过将实例锚定在由数十亿语义原型定义的全局一致潜在空间中,克服了现有嵌入方案的随机性和粒度瓶颈;其次,将生成重排器从孤立点对评估演进为对比校准的列表级策略,融合分块对比推理与校准后的绝对相关性评分,实现细粒度判别能力,同时避免传统重排方法带来的高延迟。在阿里巴巴电商平台的离线基准测试和在线A/B实验均证实,Pailitao-VL达到业界领先性能,并带来显著业务影响。本工作展示了在大规模生产环境中部署先进多模态大模型检索架构的稳健且可扩展路径。

原文摘要 · Abstract (English)

In this work, we presented Pailitao-VL, a comprehensive multi-modal retrieval system engineered for high-precision, real-time industrial search. We here address three critical challenges in the current SOTA solution: insufficient retrieval granularity, vulnerability to environmental noise, and prohibitive efficiency-performance gap. Our primary contribution lies in two fundamental paradigm shifts. First, we transitioned the embedding paradigm from traditional contrastive learning to an absolute ID-recognition task. Through anchoring instances to a globally consistent latent space defined by billions of semantic prototypes, we successfully overcome the stochasticity and granularity bottlenecks inherent in existing embedding solutions. Second, we evolved the generative reranker from isolated pointwise evaluation to the compare-and-calibrate listwise policy. By synergizing chunk-based comparative reasoning with calibrated absolute relevance scoring, the system achieves nuanced discriminative resolution while circumventing the prohibitive latency typically associated with conventional reranking methods. Extensive offline benchmarks and online A/B tests on Alibaba e-commerce platform confirm that Pailitao-VL achieves state-of-the-art performance and delivers substantial business impact. This work demonstrates a robust and scalable path for deploying advanced MLLM-based retrieval architectures in demanding, large-scale production environments.

多模态检索实时搜索电商应用重排优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。