用扩散模型生成反事实数据,提升推荐系统对冷门商品的推荐能力。
Guided Diffusion-based Counterfactual Augmentation for Robust Session-based Recommendation
- 基于扩散模型生成反事实会话数据,引导推荐模型学习非流行偏好。
- 在真实与模拟数据上,冷门商品召回率提升20%,点击率提升13%。
- 适合关注推荐公平性、减少流行度偏差的研究者和工程师。
会话推荐(SR)模型旨在根据用户当前会话中的行为推荐前K个物品。尽管已有多种SR模型被提出,但其对训练数据中固有的偏见(如流行度偏差)仍敏感,导致在真实场景下分布外数据上的性能下降。缓解流行度偏差的一种方法是反事实数据增强。与以往依赖SR模型生成数据的工作不同,本文利用先进的扩散模型生成反事实数据,提出一种基于引导扩散的反事实增强框架。通过在真实世界和模拟数据集上的离线与在线实验,结果表明该方法显著优于基线SR模型及其他先进增强框架。更重要的是,该框架在低热度目标物品上的表现显著提升,在真实与模拟数据集上分别实现召回率最高20%、点击率最高13%的增益。
原文摘要 · Abstract (English)
Session-based recommendation (SR) models aim to recommend top-K items to a user, based on the user's behaviour during the current session. Several SR models are proposed in the literature, however,concerns have been raised about their susceptibility to inherent biases in the training data (observed data) such as popularity bias. SR models when trained on the biased training data may encounter performance challenges on out-of-distribution data in real-world scenarios. One way to mitigate popularity bias is counterfactual data augmentation. Compared to prior works that rely on generating data using SR models, we focus on utilizing the capabilities of state-of-the art diffusion models for generating counterfactual data. We propose a guided diffusion-based counterfactual augmentation framework for SR. Through a combination of offline and online experiments on a real-world and simulated dataset, respectively, we show that our approach performs significantly better than the baseline SR models and other state-of-the art augmentation frameworks. More importantly, our framework shows significant improvement on less popular target items, by achieving up to 20% gain in Recall and 13% gain in CTR on real-world and simulated datasets,respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。