研究冷启动阈值如何影响推荐系统评估的可靠性
Recommendation Is a Dish Better Served Warm
- 通过逐步调整交互次数测试冷启动边界
- 发现不同阈值导致数据误删或误判,引入噪声
- 提醒研究者谨慎选择冷启动标准,提升结果可比性
现代推荐系统通常基于最小交互次数过滤冷用户和冷物品,但该阈值常被随意设定且在不同研究间差异显著,严重影响评估结果的可比性和可靠性。本文系统探究冷启动边界,通过在训练中逐步改变物品交互次数、推理中渐进更新用户历史长度,考察多个主流数据集及多种经典推荐模型的表现。结果表明,不一致的冷启动阈值可能导致有价值数据被无谓剔除,或错误将冷实例归类为热实例,从而增加系统噪声。
原文摘要 · Abstract (English)
In modern recommender systems, experimental settings typically include filtering out cold users and items based on a minimum interaction threshold. However, these thresholds are often chosen arbitrarily and vary widely across studies, leading to inconsistencies that can significantly affect the comparability and reliability of evaluation results. In this paper, we systematically explore the cold-start boundary by examining the criteria used to determine whether a user or an item should be considered cold. Our experiments incrementally vary the number of interactions for different items during training, and gradually update the length of user interaction histories during inference. We investigate the thresholds across several widely used datasets, commonly represented in recent papers from top-tier conferences, and on multiple established recommender baselines. Our findings show that inconsistent selection of cold-start thresholds can either result in the unnecessary removal of valuable data or lead to the misclassification of cold instances as warm, introducing more noise into the system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。