arXiv:2409.05461cs.IR2024-09中稿 · presentation at th…被引 5

为隐式反馈数据集的推荐算法选型提供首个系统评估,精准预测最优算法。

Recommender Systems Algorithm Selection for Ranking Prediction on Implicit Feedback Datasets

  • 构建72个数据集上24种算法的性能数据,训练多种元模型进行算法选型预测。
  • 最优元模型在未知数据集上识别最佳算法的召回率达48.6%。
  • 传统机器学习元模型(如XGBoost)优于自动化模型(如AutoGluon),效果更稳定。

针对隐式反馈数据集上的推荐系统算法选型问题,现有研究严重不足。传统方法多聚焦于显式反馈下的评分预测,而对隐式反馈下的排序预测缺乏系统探索。算法选型是推荐系统实践中的核心挑战。本文首次系统开展该方向研究:在72个推荐系统数据集上,评估24种算法在两种超参数配置下的NDCG@10表现;基于所得元数据,训练四种优化的机器学习元模型及一种自动化机器学习元模型,共三种设置。结果表明,所有元模型预测与真实排名的中位斯皮尔曼相关系数达0.857至0.918。当元模型优化为预测算法排序而非性能时,中位相关系数平均提升0.124。在预测未知数据集上最佳算法方面,最优传统元模型(如XGBoost)召回率为48.6%,优于最优自动化元模型(如AutoGluon)的47.2%。

原文摘要 · Abstract (English)

The recommender systems algorithm selection problem for ranking prediction on implicit feedback datasets is under-explored. Traditional approaches in recommender systems algorithm selection focus predominantly on rating prediction on explicit feedback datasets, leaving a research gap for ranking prediction on implicit feedback datasets. Algorithm selection is a critical challenge for nearly every practitioner in recommender systems. In this work, we take the first steps toward addressing this research gap. We evaluate the NDCG@10 of 24 recommender systems algorithms, each with two hyperparameter configurations, on 72 recommender systems datasets. We train four optimized machine-learning meta-models and one automated machine-learning meta-model with three different settings on the resulting meta-dataset. Our results show that the predictions of all tested meta-models exhibit a median Spearman correlation ranging from 0.857 to 0.918 with the ground truth. We show that the median Spearman correlation between meta-model predictions and the ground truth increases by an average of 0.124 when the meta-model is optimized to predict the ranking of algorithms instead of their performance. Furthermore, in terms of predicting the best algorithm for an unknown dataset, we demonstrate that the best optimized traditional meta-model, e.g., XGBoost, achieves a recall of 48.6%, outperforming the best tested automated machine learning meta-model, e.g., AutoGluon, which achieves a recall of 47.2%.

算法选型隐式反馈元学习推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。