通过降采样优化推荐系统能效,显著降低碳排放且保持性能。
Optimal Dataset Size for Recommender Systems: Evaluating Algorithms' Performance via Downsampling
- 采用两种降采样策略,在7个数据集上测试12种算法。
- 30%数据量可减少52%运行时间,碳排放降低至51.02 KgCO2e。
- 部分算法仅用50%数据即可保持81%性能,适合节能场景。
本研究探讨降采样在绿色推荐系统中的应用,以平衡能效与推荐质量。针对日益增长的数据规模带来的计算与环境挑战,研究对7个数据集、12种算法及两级核心剪枝配置,采用两种降采样方法进行实验。结果表明,30%的降采样比例可使单算法单数据集训练时长减少52%,碳排放降低至51.02 KgCO2e。算法性能受数据集特性、算法复杂度及降采样配置影响,部分算法在较低数据量下表现更稳定,平均仅使用50%训练集即可维持81%的全量性能;在特定配置下(固定测试集,逐步增加用户),其nDCG@10甚至超过全量数据表现。研究验证了可持续性与有效性并行的可行性,为设计节能推荐系统提供支持。
原文摘要 · Abstract (English)
This thesis investigates dataset downsampling as a strategy to optimize energy efficiency in recommender systems while maintaining competitive performance. With increasing dataset sizes posing computational and environmental challenges, this study explores the trade-offs between energy efficiency and recommendation quality in Green Recommender Systems, which aim to reduce environmental impact. By applying two downsampling approaches to seven datasets, 12 algorithms, and two levels of core pruning, the research demonstrates significant reductions in runtime and carbon emissions. For example, a 30% downsampling portion can reduce runtime by 52% compared to the full dataset, leading to a carbon emission reduction of up to 51.02 KgCO2e during the training of a single algorithm on a single dataset. The analysis reveals that algorithm performance under different downsampling portions depends on factors like dataset characteristics, algorithm complexity, and the specific downsampling configuration (scenario dependent). Some algorithms, which showed lower nDCG@10 scores compared to higher-performing ones, exhibited lower sensitivity to the amount of training data, offering greater potential for efficiency in lower downsampling portions. On average, these algorithms retained 81% of full-size performance using only 50% of the training set. In certain downsampling configurations, where more users were progressively included while keeping the test set size fixed, they even showed higher nDCG@10 scores than when using the full dataset. These findings highlight the feasibility of balancing sustainability and effectiveness, providing insights for designing energy-efficient recommender systems and promoting sustainable AI practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。