arXiv:2506.00450cs.IRcs.LG2025-06KDD被引 14

用低成本方法建模7万条用户历史,提升推荐系统表现

DV365: Extremely Long User History Modeling at Instagram

  • 通过多切片总结策略生成长序列用户嵌入
  • 支持最长7万条、平均4万条历史记录,效果优于现有模型
  • 已在Instagram和Threads落地,持续运行超一年

长用户历史是推荐系统的重要信号,但传统端到端序列建模成本高昂。本文采用离线嵌入方案,提出多切片与摘要学习策略,构建可扩展的用户表征,最大历史长度达70,000,平均40,000。该嵌入命名为DV365,已集成至Instagram和Threads的15个模型中,在先进注意力模型基础上显著提升性能,经一年多生产环境验证,具备高实用性与稳定性。

原文摘要 · Abstract (English)

Long user history is highly valuable signal for recommendation systems, but effectively incorporating it often comes with high cost in terms of data center power consumption and GPU. In this work, we chose offline embedding over end-to-end sequence length optimization methods to enable extremely long user sequence modeling as a cost-effective solution, and propose a new user embedding learning strategy, multi-slicing and summarization, that generates highly generalizable user representation of user's long-term stable interest. History length we encoded in this embedding is up to 70,000 and on average 40,000. This embedding, named as DV365, is proven highly incremental on top of advanced attentive user sequence models deployed in Instagram. Produced by a single upstream foundational model, it is launched in 15 different models across Instagram and Threads with significant impact, and has been production battle-proven for >1 year since our first launch.

推荐系统长序列建模用户画像嵌入学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。