九篇扩散模型推荐论文仅25%可复现,基线过弱导致进步假象。
Diffusion Recommender Models and the Illusion of Progress: A Concerning Study of Reproducibility and a Conceptual Mismatch
- 复现九篇顶会扩散推荐模型,发现仅四分之一结果可重复。
- 调优后简单基线始终优于原文报告的扩散模型效果。
- 扩散模型与推荐任务特性不匹配,生成能力也被严重限制。
每年有大量新机器学习模型被发表,宣称显著提升推荐系统的性能。然而早期可复现性研究指出,由于广泛存在的方法学问题(如未调优的基线对比),实际进展可能有限,造成进步的假象。本文针对近年来快速发展的去噪扩散概率模型在推荐系统中的应用,尝试复现九篇来自SIGIR 2023和2024的扩散推荐算法。结果显示,仅有25%的报告结果可完全复现;且因原始论文使用了弱基线,无法证明扩散模型优于现有先进方法。在受控评估中,经过充分调优的简单基线模型始终优于原文报道的扩散模型表现。此外,我们识别出扩散模型特性与传统top-n推荐任务需求之间存在关键错配,质疑其在推荐场景中的适用性。同时,这些论文中模型的生成能力被严重限制至最低水平。整体结果呼吁该领域加强科学严谨性,并推动研究与发表文化的根本性变革。
原文摘要 · Abstract (English)
Countless new machine learning models are published every year and are reported to significantly advance the state-of-the-art in top-n recommendation. However, earlier reproducibility studies indicate that progress in this area may be quite limited, due to widespread methodological issues, e.g., comparisons with untuned baseline models, creating an illusion of progress. In this work, we examine whether these problems persist in today's research by attempting to reproduce nine SIGIR 2023 and 2024 recommendation algorithms based on Denoising Diffusion Probabilistic Models, a recent but rapidly expanding research area. Only 25% of reported results are fully reproducible and, since the original papers relied on weak baselines, they do not establish the superiority of diffusion models over state-of-the-art methods. In our controlled evaluations, well-tuned simpler baselines consistently exceed the diffusion-based models' effectiveness reported in the original papers. Furthermore, we identify key mismatches between the characteristics of diffusion models and those of the traditional top-n recommendation task, raising doubts about their suitability for recommendation. Moreover, in the analyzed papers, the generative capabilities of these models are constrained to a minimum. Overall, our results call for greater scientific rigor and a disruptive change in the research and publication culture in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。