评测大模型在推荐系统中的认知能力,发现其擅长浅层关联但难捕捉用户深层偏好。
The Mental World of Large Language Models in Recommendation: A Benchmark on Association, Personalization, and Knowledgeability
- 构建38K样本的LRWorld基准,从关联、个性化、知识性三维度评估大模型
- 大模型在物品相似度和实体关系推理上表现好,但对深度行为嵌入理解不足
- 适合关注大模型在推荐中局限性的研究者,尤其关注多模态与噪声鲁棒性
大型语言模型(LLMs)在推荐系统(RecSys)中展现出潜力,可作为知识增强器或零样本排序器。然而,其内在的语言世界知识与推荐系统的个性化行为世界之间存在显著语义鸿沟。当前研究缺乏全面的基准来评估大模型在推荐系统中的局限与边界。为此,我们提出名为LRWorld的基准,包含超过38,000个高质量样本和2300万词元,由多个公开推荐数据集精心构建生成。该基准将大模型在推荐系统中的“心智世界”分为关联、个性化和知识性三个尺度,涵盖十项因素与31项任务。基于此,对数十个大模型的综合实验表明,它们尚未有效捕捉深层神经个性化嵌入,但在浅层记忆型物品-物品相似度上表现良好;在推断用户兴趣时,能较好感知物品实体关系、层级分类体系及物品关联规则。此外,大模型在多模态知识推理(如电影海报、产品图像)和噪声用户画像鲁棒性方面表现出色。但无一模型在全部十项因素上表现一致优异。进一步分析了模型规模、位置偏差等因素的影响。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown potential in recommendation systems (RecSys) by using them as either knowledge enhancer or zero-shot ranker. A key challenge lies in the large semantic gap between LLMs and RecSys where the former internalizes language world knowledge while the latter captures personalized world of behaviors. Unfortunately, the research community lacks a comprehensive benchmark that evaluates the LLMs over their limitations and boundaries in RecSys so that we can draw a confident conclusion. To investigate this, we propose a benchmark named LRWorld containing over 38K high-quality samples and 23M tokens carefully compiled and generated from widely used public recommendation datasets. LRWorld categorizes the mental world of LLMs in RecSys as three main scales (association, personalization, and knowledgeability) spanned by ten factors with 31 measures (tasks). Based on LRWorld, comprehensive experiments on dozens of LLMs show that they are still not well capturing the deep neural personalized embeddings but can achieve good results on shallow memory-based item-item similarity. They are also good at perceiving item entity relations, entity hierarchical taxonomies, and item-item association rules when inferring user interests. Furthermore, LLMs show a promising ability in multimodal knowledge reasoning (movie poster and product image) and robustness to noisy profiles. None of them show consistently good performance over the ten factors. Model sizes, position bias, and more are ablated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。