融合协同与语义信息,提升推荐系统对用户全场景兴趣的捕捉能力
PRECISE: Pre-training Sequential Recommenders with Collaborative and Semantic Information
- 通过协同信号与语义信息联合建模,构建统一用户兴趣表征
- 在公开与工业数据集上均显著优于现有方法,尤其在长尾和冷启动场景表现突出
- 适合需要跨场景推荐、关注用户长期兴趣建模的研究者与工程师
现实中的推荐系统常为用户提供多样化的交互场景。由于工业平台用户数量庞大,难以用单一统一模型满足所有场景需求,通常为每个独立场景建立单独推荐管道,这导致难以全面把握用户兴趣。近期研究尝试通过预训练模型来整合用户整体兴趣。传统预训练推荐模型主要依赖协同信号,但难以处理长尾物品和冷启动问题。随着大语言模型的发展,利用其提取用户与物品语义信息的研究日益增多,但文本推荐高度依赖复杂特征工程,且常无法捕捉协同相似性。为此,我们提出一种新的序列推荐预训练框架PRECISE,融合协同信号与语义信息,并采用分阶段学习机制:先建模跨场景的全局兴趣,再聚焦目标场景的特定行为兴趣。实证表明,PRECISE能精准捕捉用户全范围兴趣,并有效迁移至目标场景。实验结果证明,该框架在公共与工业数据集上均取得优异性能。
原文摘要 · Abstract (English)
Real-world recommendation systems commonly offer diverse content scenarios for users to interact with. Considering the enormous number of users in industrial platforms, it is infeasible to utilize a single unified recommendation model to meet the requirements of all scenarios. Usually, separate recommendation pipelines are established for each distinct scenario. This practice leads to challenges in comprehensively grasping users' interests. Recent research endeavors have been made to tackle this problem by pre-training models to encapsulate the overall interests of users. Traditional pre-trained recommendation models mainly capture user interests by leveraging collaborative signals. Nevertheless, a prevalent drawback of these systems is their incapacity to handle long-tail items and cold-start scenarios. With the recent advent of large language models, there has been a significant increase in research efforts focused on exploiting LLMs to extract semantic information for users and items. However, text-based recommendations highly rely on elaborate feature engineering and frequently fail to capture collaborative similarities. To overcome these limitations, we propose a novel pre-training framework for sequential recommendation, termed PRECISE. This framework combines collaborative signals with semantic information. Moreover, PRECISE employs a learning framework that initially models users' comprehensive interests across all recommendation scenarios and subsequently concentrates on the specific interests of target-scene behaviors. We demonstrate that PRECISE precisely captures the entire range of user interests and effectively transfers them to the target interests. Empirical findings reveal that the PRECISE framework attains outstanding performance on both public and industrial datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。