构建首个学术文献推荐的自然语言用户画像数据集,助力可解释推荐研究。
SciNUP: Natural Language User Interest Profiles for Scientific Literature Recommendation
- 基于作者发表记录生成自然语言用户画像,模拟真实阅读偏好。
- 多类推荐方法表现相当但召回结果差异大,体现互补性。
- 适合关注可解释推荐、个性化检索的研究者使用。
在推荐系统中使用自然语言(NL)用户画像相比传统表示方式更具透明度和用户可控性。然而,目前缺乏大规模、公开可用的测试集来评估基于自然语言画像的推荐效果。为填补这一空白,我们提出SciNUP,一个用于学术文献推荐的新颖合成数据集,利用作者的出版历史生成自然语言画像及对应的真值推荐项。我们基于该数据集对多种基线方法进行了比较,涵盖稀疏与稠密检索方法,以及最先进的大模型重排序器。结果显示,尽管各类方法性能相近,但它们常推荐不同项目,表明其行为具有互补性。同时,仍存在显著提升空间,凸显高效自然语言推荐方法的必要性。因此,SciNUP为该领域的未来研究与开发提供了宝贵资源。
原文摘要 · Abstract (English)
The use of natural language (NL) user profiles in recommender systems offers greater transparency and user control compared to traditional representations. However, there is scarcity of large-scale, publicly available test collections for evaluating NL profile-based recommendation. To address this gap, we introduce SciNUP, a novel synthetic dataset for scholarly recommendation that leverages authors' publication histories to generate NL profiles and corresponding ground truth items. We use this dataset to conduct a comparison of baseline methods, ranging from sparse and dense retrieval approaches to state-of-the-art LLM-based rerankers. Our results show that while baseline methods achieve comparable performance, they often retrieve different items, indicating complementary behaviors. At the same time, considerable headroom for improvement remains, highlighting the need for effective NL-based recommendation approaches. The SciNUP dataset thus serves as a valuable resource for fostering future research and development in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。