推荐系统评估中单次随机种子可能误导稳定性结论,需重视种子影响。
Training seeds and model-selection stability in recommender-system evaluation

- 固定数据划分,变换训练种子,检验其对评估结果的影响
- 种子变化可导致用户指标、模型选择和推荐列表差异显著
- 建议将种子纳入评估协议,而非视为无关噪声
推荐系统实验常依赖单一随机训练种子,假设运行间随机性对评估结论影响有限。这一假设存在风险,因为训练种子会影响参数初始化、小批量顺序、丢弃率、掩码、潜在变量采样及训练时负采样等算法相关机制。本文在固定数据划分的前提下,通过改变训练种子来分析其在不同超参数配置下的影响,从用户级指标敏感性、基于验证的模型选择以及推荐列表一致性三个层面进行考察。结果表明,种子变化通常可被检测到,其影响取决于配置是否清晰分离、验证结果能否迁移至测试集,以及相似得分是否产生相似的Top-k推荐列表。研究提示:仅报告单种子结果可能过度高估推荐系统评估的稳定性,训练种子应作为评估协议的一部分,而非偶然实现噪声。
原文摘要 · Abstract (English)
Recommender-system experiments often rely on a single random training seed, assuming that run-to-run stochasticity has limited impact on evaluation conclusions. This assumption is risky, as a training seed may influence several algorithm-dependent mechanisms, including parameter initialization, mini-batch ordering, dropout, masking, latent sampling, and training-time negative sampling. We examine this assumption by fixing the data partition and varying the training seed across hyperparameter configurations. We analyze seed effects at three levels: user-level metric sensitivity, validation-based model selection and recommendation-list agreement. Results show that seed variation is often detectable. Its impact depends on whether configurations are clearly separated, whether validation results transfer to test, and whether similar scores lead to similar top-$k$ lists. Findings suggest that reporting single-seed results can overstate the stability of recommender system evaluation, and that training seeds should be treated as part of the evaluation protocol rather than as incidental implementation noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。