arXiv:2501.08695cs.IR2025-01KDD被引 16

实时向量量化索引提升大规模推荐系统召回效率

Real-time Indexing for Large-scale Recommendation by Streaming Vector Quantization Retriever

  • 提出流式向量量化索引,支持实时加索引
  • 在抖音等平台部署后用户参与度显著提升
  • 兼顾复杂排序模型,适合工业级推荐系统

检索器是推荐系统中关键环节,在严格延迟限制下需高效筛选潜在正样本。现有大规模系统多依赖近似计算和索引粗略缩小候选集,采用简单排序模型。由于简单模型预测精度不足,多数方法转向引入复杂排序模型,但索引有效性这一根本问题仍未解决,制约了模型复杂度提升。本文提出新型索引结构——流式向量量化(Streaming VQ),作为新一代检索范式。该结构支持实时为物品添加索引,具备即时性;通过细致验证多种变体,还实现了索引均衡与可修复性,能兼容复杂排序模型。作为轻量且易实现的架构,Streaming VQ已在抖音及抖音极速版全面部署,替代所有主要检索器,带来显著用户参与度提升。

原文摘要 · Abstract (English)

Retrievers, which form one of the most important recommendation stages, are responsible for efficiently selecting possible positive samples to the later stages under strict latency limitations. Because of this, large-scale systems always rely on approximate calculations and indexes to roughly shrink candidate scale, with a simple ranking model. Considering simple models lack the ability to produce precise predictions, most of the existing methods mainly focus on incorporating complicated ranking models. However, another fundamental problem of index effectiveness remains unresolved, which also bottlenecks complication. In this paper, we propose a novel index structure: streaming Vector Quantization model, as a new generation of retrieval paradigm. Streaming VQ attaches items with indexes in real time, granting it immediacy. Moreover, through meticulous verification of possible variants, it achieves additional benefits like index balancing and reparability, enabling it to support complicated ranking models as existing approaches. As a lightweight and implementation-friendly architecture, streaming VQ has been deployed and replaced all major retrievers in Douyin and Douyin Lite, resulting in remarkable user engagement gain.

推荐系统向量量化实时索引

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。