arXiv:2511.11649cs.IRcs.LG2025-11被引 1

对比推荐系统集成方法的能耗,发现选优策略更省电

The Environmental Impact of Ensemble Techniques in Recommender Systems

  • 用智能插座实测不同集成策略的耗电,比较准确率与能耗
  • 选优集成仅多耗18.8%电就提升0.96%准确率,比平均法省电88%以上
  • 在超大规模数据上,集成方法能耗暴涨20倍,碳排放达单模型20倍

推荐系统中的集成技术虽能提升准确率10-30%,但其环境影响尚未评估。本研究在Surprise和LensKit框架下,于四个涵盖10万至780万交互的数据集上开展93次实验,对比四种集成策略(平均、加权、堆叠/排序融合、前几佳)与单一优化模型的性能与能耗。使用智能插座测量能耗,结果揭示准确率与能耗呈非线性关系:集成方法带来0.3%-5.7%的准确率提升,但能耗增加19%-2549%。其中,'前几佳'集成策略效率最高,在MovieLens-1M上仅增18.8%能耗即获0.96% RMSE改善;在MovieLens-100K上实现5.7% NDCG提升,能耗增加103%。而全平均策略能耗高出88%-270%却收益相近。在最大数据集Anime(780万交互)上,Surprise集成耗电达0.21瓦时(单模型0.01瓦时),能耗增长2005%,仅提升1.2%准确率,碳排放从2.6毫克增至53.8毫克。研究首次系统量化集成推荐系统的能效与碳足迹,表明选择性集成优于全量平均,且在工业级规模下存在显著扩展瓶颈。

原文摘要 · Abstract (English)

Ensemble techniques in recommender systems have demonstrated accuracy improvements of 10-30%, yet their environmental impact remains unmeasured. While deep learning recommendation algorithms can generate up to 3,297 kg CO2 per paper, ensemble methods have not been sufficiently evaluated for energy consumption. This thesis investigates how ensemble techniques influence environmental impact compared to single optimized models. We conducted 93 experiments across two frameworks (Surprise for rating prediction, LensKit for ranking) on four datasets spanning 100,000 to 7.8 million interactions. We evaluated four ensemble strategies (Average, Weighted, Stacking/Rank Fusion, Top Performers) against simple baselines and optimized single models, measuring energy consumption with a smart plug. Results revealed a non-linear accuracy-energy relationship. Ensemble methods achieved 0.3-5.7% accuracy improvements while consuming 19-2,549% more energy depending on dataset size and strategy. The Top Performers ensemble showed best efficiency: 0.96% RMSE improvement with 18.8% energy overhead on MovieLens-1M, and 5.7% NDCG improvement with 103% overhead on MovieLens-100K. Exhaustive averaging strategies consumed 88-270% more energy for comparable gains. On the largest dataset (Anime, 7.8M interactions), the Surprise ensemble consumed 2,005% more energy (0.21 Wh vs. 0.01 Wh) for 1.2% accuracy improvement, producing 53.8 mg CO2 versus 2.6 mg CO2 for the single model. This research provides one of the first systematic measurements of energy and carbon footprint for ensemble recommender systems, demonstrates that selective strategies offer superior efficiency over exhaustive averaging, and identifies scalability limitations at industrial scale. These findings enable informed decisions about sustainable algorithm selection in recommender systems.

推荐系统能耗评估可持续计算集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。