五款热门意图感知推荐模型无法复现,且均被传统方法超越。
A Worrying Reproducibility Study of Intent-Aware Recommendation Models
- 复现五个顶级论文的意图感知推荐模型
- 所有新模型均被传统非神经方法超越
- 揭示推荐系统研究中的可复现性危机
近期,意图感知推荐系统(IARS)受到广泛关注,其核心理念是通过预测用户潜在动机和短期目标来提升推荐效果。为此,多个复杂神经模型被提出。然而,在复杂神经推荐系统领域,已有大量研究指出复现困难且性能提升可能源于弱基线对比。本文针对这一问题,尝试复现五篇发表于顶级会议的IARS模型,并与多种传统非神经推荐模型进行对比。结果显示:在两例中,使用论文提供的最优超参数运行代码也未能复现原文结果;更令人担忧的是,所有被检视的IARS方法均被至少一种传统模型超越。这些发现揭示了该领域持续存在的方法论缺陷,亟需更严谨的学术实践。
原文摘要 · Abstract (English)
Lately, we have observed a growing interest in intent-aware recommender systems (IARS). The promise of such systems is that they are capable of generating better recommendations by predicting and considering the underlying motivations and short-term goals of consumers. From a technical perspective, various sophisticated neural models were recently proposed in this emerging and promising area. In the broader context of complex neural recommendation models, a growing number of research works unfortunately indicates that (i) reproducing such works is often difficult and (ii) that the true benefits of such models may be limited in reality, e.g., because the reported improvements were obtained through comparisons with untuned or weak baselines. In this work, we investigate if recent research in IARS is similarly affected by such problems. Specifically, we tried to reproduce five contemporary IARS models that were published in top-level outlets, and we benchmarked them against a number of traditional non-neural recommendation models. In two of the cases, running the provided code with the optimal hyperparameters reported in the paper did not yield the results reported in the paper. Worryingly, we find that all examined IARS approaches are consistently outperformed by at least one traditional model. These findings point to sustained methodological issues and to a pressing need for more rigorous scholarly practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。