arXiv:2505.14241cs.IR2025-05被引 1

用采样加速图推荐系统训练,50%数据仍保性能,但过少数据性能骤降。

The Limits of Graph Samplers for Training Inductive Recommender Systems: Extended results

  • 在真实数据集上测试六种采样方法,验证子图泛化可行性。
  • 仅用50%训练数据可降低86%训练时间,性能基本不变。
  • 采样需考虑时间维度,未来需设计新采样与推荐方法。

归纳式推荐系统能为新用户和新物品推荐,无需重新训练。然而,这些方法仍需使用全部数据进行训练,单次模型训练需数天,不计超参数调优时间。本文聚焦基于图的推荐系统,即建模为异质网络的系统。在其他应用中,图采样可通过子图研究推断原图特性。因此,我们考察采样技术在此任务中的适用性。在三个真实数据集上,采用三种先进归纳式方法及六种采样策略进行测试。结果表明,使用50%训练数据可实现高达86%的训练时间减少,且性能保持稳定;但进一步减少数据会导致性能显著下降。此外,针对推荐数据,图采样还需考虑时间维度。因此,若需更高数据压缩率,应研究新型图采样方法,并设计新的归纳式推荐模型。

原文摘要 · Abstract (English)

Inductive Recommender Systems are capable of recommending for new users and with new items thus avoiding the need to retrain after new data reaches the system. However, these methods are still trained on all the data available, requiring multiple days to train a single model, without counting hyperparameter tuning. In this work we focus on graph-based recommender systems, i.e., systems that model the data as a heterogeneous network. In other applications, graph sampling allows to study a subgraph and generalize the findings to the original graph. Thus, we investigate the applicability of sampling techniques for this task. We test on three real world datasets, with three state-of-the-art inductive methods, and using six different sampling methods. We find that its possible to maintain performance using only 50% of the training data with up to 86% percent decrease in training time; however, using less training data leads to far worse performance. Further, we find that when it comes to data for recommendations, graph sampling should also account for the temporal dimension. Therefore, we find that if higher data reduction is needed, new graph based sampling techniques should be studied and new inductive methods should be designed.

图推荐采样训练加速归纳学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。