用可动态演化的兴趣空间提升推荐系统发现新兴趣的能力
SPARC: Soft Probabilistic Adaptive multi-interest Retrieval Model via Codebooks for recommender system
- 构建离散兴趣空间,让兴趣随用户行为实时演化
- 在线推理采用概率软搜索,使推荐从被动匹配转为主动探索
- 在千万级平台测试中显著提升内容曝光与用户停留时长
建模多兴趣已成为现实推荐系统的核心问题。现有方法存在三大挑战:1)兴趣通常来自预定义外部知识,固定不变,无法随用户实时偏好动态演化;2)在线推理普遍采用过度利用策略,仅匹配已有兴趣,缺乏对新颖和长尾兴趣的主动探索。为此,我们提出新型检索框架 SPARC(Soft Probabilistic Adaptive Retrieval Model via Codebooks)。首先,通过残差量化变分自编码器(RQ-VAE)构建离散兴趣空间,联合训练工业级大规模推荐模型,挖掘感知用户反馈、可动态演化的行为感知兴趣。其次,设计概率兴趣模块,预测整个动态离散兴趣空间上的概率分布,支持在线推理中的高效“软搜索”策略,将检索范式从“被动匹配”转变为“主动探索”,有效促进兴趣发现。在拥有数千万日活用户的工业平台进行在线 A/B 测试,业务指标显著提升:用户观看时长增加 0.9%,页面浏览量(PV)提升 0.4%,新内容 24 小时内达到 500 次浏览的占比提升 22.7%。离线实验在开源 Amazon Product 数据集上,召回率(Recall@K)与归一化折现累计增益(NDCG@K)也一致提升。线上线下实验共同验证了该方法的有效性与实际价值。
原文摘要 · Abstract (English)
Modeling multi-interests has arisen as a core problem in real-world RS. Current multi-interest retrieval methods pose three major challenges: 1) Interests, typically extracted from predefined external knowledge, are invariant. Failed to dynamically evolve with users' real-time consumption preferences. 2) Online inference typically employs an over-exploited strategy, mainly matching users' existing interests, lacking proactive exploration and discovery of novel and long-tail interests. To address these challenges, we propose a novel retrieval framework named SPARC(Soft Probabilistic Adaptive Retrieval Model via Codebooks). Our contribution is two folds. First, the framework utilizes Residual Quantized Variational Autoencoder (RQ-VAE) to construct a discretized interest space. It achieves joint training of the RQ-VAE with the industrial large scale recommendation model, mining behavior-aware interests that can perceive user feedback and evolve dynamically. Secondly, a probabilistic interest module that predicts the probability distribution over the entire dynamic and discrete interest space. This facilitates an efficient "soft-search" strategy during online inference, revolutionizing the retrieval paradigm from "passive matching" to "proactive exploration" and thereby effectively promoting interest discovery. Online A/B tests on an industrial platform with tens of millions daily active users, have achieved substantial gains in business metrics: +0.9% increase in user view duration, +0.4% increase in user page views (PV), and a +22.7% improvement in PV500(new content reaching 500 PVs in 24 hours). Offline evaluations are conducted on open-source Amazon Product datasets. Metrics, such as Recall@K and Normalized Discounted Cumulative Gain@K(NDCG@K), also showed consistent improvement. Both online and offline experiments validate the efficacy and practical value of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。