arXiv:2503.11924cs.CLcs.AI2025-03被引 1

构建含用户批评与叙事的推荐数据集,提升大模型对话推荐能力

REGEN: A Dataset and Benchmarks with Natural Language Critiques and Narratives

  • 基于亚马逊评论扩展生成用户批评和物品叙事文本
  • 引入新任务框架,联合建模推荐与上下文叙述生成
  • 适合研究对话推荐、大模型应用与多任务学习的学者

本文提出新型数据集REGEN(Reviews Enhanced with GEnerative Narratives),用于评估推荐型大语言模型的对话能力。该数据集在亚马逊产品评论基础上,通过补全两个关键自然语言特征:一是用户批评(代表引导后续选择的查询),二是与推荐项相关的丰富叙事文本,包括产品推荐语、购买解释及偏好总结。同时,我们建立端到端建模范式,要求模型根据用户历史(物品与批评)生成推荐结果及对应叙事。为此提出统一多任务框架LUMEN,以大语言模型为骨干,实现批评、检索与生成一体化。通过自动评分评估数据质量,并训练传统与大模型推荐系统进行对比。实验表明,引入批评可显著提升推荐质量,使模型融合语言理解与推荐信号;使用该数据集训练的大模型能有效生成推荐与上下文叙事,性能达到当前先进水平。

原文摘要 · Abstract (English)

This paper introduces a novel dataset REGEN (Reviews Enhanced with GEnerative Narratives), designed to benchmark the conversational capabilities of recommender Large Language Models (LLMs), addressing the limitations of existing datasets that primarily focus on sequential item prediction. REGEN extends the Amazon Product Reviews dataset by inpainting two key natural language features: (1) user critiques, representing user "steering" queries that lead to the selection of a subsequent item, and (2) narratives, rich textual outputs associated with each recommended item taking into account prior context. The narratives include product endorsements, purchase explanations, and summaries of user preferences. Further, we establish an end-to-end modeling benchmark for the task of conversational recommendation, where models are trained to generate both recommendations and corresponding narratives conditioned on user history (items and critiques). For this joint task, we introduce a modeling framework LUMEN (LLM-based Unified Multi-task Model with Critiques, Recommendations, and Narratives) which uses an LLM as a backbone for critiquing, retrieval and generation. We also evaluate the dataset's quality using standard auto-rating techniques and benchmark it by training both traditional and LLM-based recommender models. Our results demonstrate that incorporating critiques enhances recommendation quality by enabling the recommender to learn language understanding and integrate it with recommendation signals. Furthermore, LLMs trained on our dataset effectively generate both recommendations and contextual narratives, achieving performance comparable to state-of-the-art recommenders and language models.

对话推荐大模型数据集自然语言生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。