研究离线推荐评估设计如何影响结果可靠性,发现无最优方案。
On the Convergent Validity of Offline Evaluation Designs for Recommender Systems
- 对比不同数据过滤与候选集构建方式对评估结果的影响
- 验证稀疏数据评估与真实反馈排序的相关性差异
- 提醒研究者根据数据集和目标选择合适评估配置
离线评估依赖历史交互日志,但数据稀疏、不完整或有偏差,引发对其是否反映真实用户偏好的质疑。本文研究评估设计选择如何影响推荐系统比较的有效性。在多个数据集上,测试不同设置下的推荐模型,包括数据过滤阈值和候选集构建方式。通过比较稀疏数据训练-测试分割得到的模型排名与基于密集真实反馈的排名相关性,评估各配置的有效性。结果表明,评估有效性取决于数据集和具体评估目标,不存在统一最优的离线评估设计。
原文摘要 · Abstract (English)
Offline evaluation on historical interaction logs is the most common evaluation methodology for recommender systems. However, such evaluations depend on sparse, incomplete, or biased data, which raises concerns about whether commonly used evaluation setups reliably reflect true user preferences. In this work, we study how offline evaluation design choices affect the validity of recommender system comparisons. We evaluate a set of recommendation models across several evaluation setups that vary key factors such as data filtering thresholds and candidate set construction. To assess the validity of these configurations, we measure the correlation between model rankings obtained from conventional train-test splits on sparse interaction data and rankings from evaluations based on dense ground-truth user feedback. We use this agreement as an indication of their validity with respect to true user preferences. Our results show that the validity of sparse evaluation depends on the dataset and the specific dense evaluation targets, and that there is no uniformly best offline evaluation design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。