arXiv:2503.09382cs.IRcs.AI2025-03中稿 · publication at WSD…被引 27

为大模型推荐助手设计新评测基准,提升真实场景评估能力

Towards Next-Generation Recommender Systems: A Benchmark for Personalized Recommendation Assistant with LLMs

  • 构建涵盖复杂需求的多难度用户查询数据集
  • 发现大模型更擅长处理明确条件,推理和误导信息应对弱
  • 适合研究智能推荐系统与大模型交互的学者使用

推荐系统广泛应用于各类数字平台,传统方法局限于固定简单场景,难以适应交互式新任务。大语言模型(LLMs)的兴起推动推荐系统向智能个性化助手演进。然而,现有研究多依赖固定提示模板,且缺乏高质量文本查询数据,限制了对大模型推荐能力的全面评估。为此,我们提出 RecBench+,一个面向 LLM 推荐助手的新基准数据集,包含涵盖硬性条件与软性偏好、不同难度级别的多样化查询。在 RecBench+ 上评估主流 LLMs 发现:1)大模型初步具备推荐助手能力;2)对显式条件处理较好,但在需推理或含误导信息的查询上表现不佳。数据集已开源:https://github.com/jiani-huang/RecBench.git。

原文摘要 · Abstract (English)

Recommender systems (RecSys) are widely used across various modern digital platforms and have garnered significant attention. Traditional recommender systems usually focus only on fixed and simple recommendation scenarios, making it difficult to generalize to new and unseen recommendation tasks in an interactive paradigm. Recently, the advancement of large language models (LLMs) has revolutionized the foundational architecture of RecSys, driving their evolution into more intelligent and interactive personalized recommendation assistants. However, most existing studies rely on fixed task-specific prompt templates to generate recommendations and evaluate the performance of personalized assistants, which limits the comprehensive assessments of their capabilities. This is because commonly used datasets lack high-quality textual user queries that reflect real-world recommendation scenarios, making them unsuitable for evaluating LLM-based personalized recommendation assistants. To address this gap, we introduce RecBench+, a new dataset benchmark designed to access LLMs' ability to handle intricate user recommendation needs in the era of LLMs. RecBench+ encompasses a diverse set of queries that span both hard conditions and soft preferences, with varying difficulty levels. We evaluated commonly used LLMs on RecBench+ and uncovered below findings: 1) LLMs demonstrate preliminary abilities to act as recommendation assistants, 2) LLMs are better at handling queries with explicitly stated conditions, while facing challenges with queries that require reasoning or contain misleading information. Our dataset has been released at https://github.com/jiani-huang/RecBench.git.

推荐系统大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。