arXiv:2503.07823cs.IRcs.DL2025-03被引 6

复现10篇推荐系统论文发现数据泄露、代码不符、基线虚高问题。

Reproducibility and Artifact Consistency of the SIGIR 2022 Recommender Systems Papers Based on Message Passing

  • 分析SIGIR 2022的10篇图神经网络推荐论文
  • 发现数据划分错误、测试集信息泄露等严重问题
  • 提醒研究者警惕复杂基线带来的虚假性能提升

基于消息传递的图神经网络方法在推荐系统中日益流行,2022与2023年SIGIR会议中均有相关论文发表。本文对10篇主要来自SIGIR 2022的图推荐系统论文进行复现分析,评估其对后续SIGIR 2023工作的影响力。结果揭示三大关键问题:(i) 存在大量不良实践,如训练与测试数据间的信息泄露、错误的数据划分,严重质疑结果有效性;(ii) 提供的源代码与数据与论文描述存在频繁不一致,导致实际评估内容不明;(iii) 倾向于使用新或复杂的基线模型,而这些基线反而弱于简单基线,尤其在Amazon-Book数据集上,最先进性能显著退化。由于上述问题,我们无法验证所考察论文中的多数结论。

原文摘要 · Abstract (English)

Graph-based techniques relying on neural networks and embeddings have gained attention as a way to develop Recommender Systems (RS) with several papers on the topic presented at SIGIR 2022 and 2023. Given the importance of ensuring that published research is methodologically sound and reproducible, in this paper we analyze 10 graph-based RS papers, most of which were published at SIGIR 2022, and assess their impact on subsequent work published in SIGIR 2023. Our analysis reveals several critical points that require attention: (i) the prevalence of bad practices, such as erroneous data splits or information leakage between training and testing data, which call into question the validity of the results; (ii) frequent inconsistencies between the provided artifacts (source code and data) and their descriptions in the paper, causing uncertainty about what is actually being evaluated; and (iii) the preference for new or complex baselines that are weaker compared to simpler ones, creating the impression of continuous improvement even when, particularly for the Amazon-Book dataset, the state-of-the-art has significantly worsened. Due to these issues, we are unable to confirm the claims made in most of the papers that we examined and attempted to reproduce.

推荐系统可复现性图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。