arXiv:2410.09359cs.IRcs.AI2024-10被引 11

通过数据采样优化推荐系统,减少能耗且不影响效果。

Green Recommender Systems: Optimizing Dataset Size for Energy-Efficient Algorithm Performance

  • 用采样法控制数据量,测试不同规模下算法表现。
  • 部分算法在数据减半时仍保持90%以上性能(如nDCG@10)。
  • 适合关注绿色计算与节能部署的推荐系统研究者。

随着推荐系统广泛应用,大规模模型训练带来的环境影响和能效问题日益受到关注。本文在绿色推荐系统的背景下,通过数据采样技术优化数据集规模,探究能效与算法性能的平衡。在MovieLens 100K、1M、10M及Amazon Toys and Games数据集上,测试了多种推荐算法在不同数据比例下的表现。结果表明,虽然更多数据通常带来更好性能,但某些算法(如FunkSVD、BiasedMF)在非均衡、稀疏数据集(如Amazon Toys and Games)中,可将训练数据减少50%,仍保持nDCG@10分数约在全量数据性能的13%以内。这说明通过策略性缩减数据规模,可在不显著牺牲推荐质量的前提下降低计算与环境成本。本研究为构建可持续、绿色的推荐系统提供了关键实践依据。

原文摘要 · Abstract (English)

As recommender systems become increasingly prevalent, the environmental impact and energy efficiency of training large-scale models have come under scrutiny. This paper investigates the potential for energy-efficient algorithm performance by optimizing dataset sizes through downsampling techniques in the context of Green Recommender Systems. We conducted experiments on the MovieLens 100K, 1M, 10M, and Amazon Toys and Games datasets, analyzing the performance of various recommender algorithms under different portions of dataset size. Our results indicate that while more training data generally leads to higher algorithm performance, certain algorithms, such as FunkSVD and BiasedMF, particularly with unbalanced and sparse datasets like Amazon Toys and Games, maintain high-quality recommendations with up to a 50% reduction in training data, achieving nDCG@10 scores within approximately 13% of full dataset performance. These findings suggest that strategic dataset reduction can decrease computational and environmental costs without substantially compromising recommendation quality. This study advances sustainable and green recommender systems by providing insights for reducing energy consumption while maintaining effectiveness.

绿色计算推荐系统数据采样能效优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。