arXiv:2411.08700cs.IRcs.AI2024-11

优化新闻推荐中的负样本采样,提升准确率并降低模型复杂度

Rethinking negative sampling in content-based news recommendation

  • 设计新型负样本采样策略,增强模型对短时效新闻的适应性
  • 在MIND数据集上达到SOTA精度,训练速度更快且模型更轻量
  • 适合关注隐私保护与可扩展性的推荐系统开发者

新闻推荐系统受限于文章生命周期短暂,存在快速相关性衰减问题。现有基于内容的神经方法虽有潜力,但常依赖复杂架构且忽视负样本选择。本文认为负样本采样对模型表现影响显著,提出一种新采样技术,在保持高精度的同时降低模型复杂度并加速训练。基于MIND数据集的实验表明,该方法性能可媲美当前最优模型。此外,该策略有助于实现推荐系统的去中心化,从而提升隐私保护与可扩展性。

原文摘要 · Abstract (English)

News recommender systems are hindered by the brief lifespan of articles, as they undergo rapid relevance decay. Recent studies have demonstrated the potential of content-based neural techniques in tackling this problem. However, these models often involve complex neural architectures and often lack consideration for negative examples. In this study, we posit that the careful sampling of negative examples has a big impact on the model's outcome. We devise a negative sampling technique that not only improves the accuracy of the model but also facilitates the decentralization of the recommendation system. The experimental results obtained using the MIND dataset demonstrate that the accuracy of the method under consideration can compete with that of State-of-the-Art models. The utilization of the sampling technique is essential in reducing model complexity and accelerating the training process, while maintaining a high level of accuracy. Finally, we discuss how decentralized models can help improve privacy and scalability.

新闻推荐负样本采样去中心化模型轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。