arXiv:2412.06949cs.IRcs.AI2024-12被引 7

用Reddit数据补全对话推荐中的用户偏好信息,提升推荐准确率。

Bridging Conversational and Collaborative Signals for Conversational Recommendation

  • 用Reddit对话与MovieLens数据关联,增强物品表征
  • 相比最佳基线,命中率提升12.32%,NDCG提升9.9%
  • 适合做对话推荐且关注协同过滤融合的研究者

对话推荐系统(CRS)依赖对话上下文生成推荐,但常因缺乏协同过滤(CF)信号而表现不佳,CF信号能捕捉用户-物品交互模式。我们构建了Reddit-ML32M数据集,将Reddit对话与MovieLens 32M的交互数据关联,利用协同知识丰富物品表征,缓解对话数据中交互稀疏问题。提出基于大语言模型(LLM)的框架,通过Reddit-ML32M对齐LLM生成的推荐与CF嵌入,优化排序结果。在三类基线(仅用CRS交互的CF推荐、传统对话推荐模型、仅依赖对话上下文的LLM方法)上评估,本方法持续提升性能,命中率提高12.32%,NDCG提升9.9%,超越依赖对话上下文但无协同物品表征的最佳基线。

原文摘要 · Abstract (English)

Conversational recommendation systems (CRS) leverage contextual information from conversations to generate recommendations but often struggle due to a lack of collaborative filtering (CF) signals, which capture user-item interaction patterns essential for accurate recommendations. We introduce Reddit-ML32M, a dataset that links Reddit conversations with interactions on MovieLens 32M, to enrich item representations by leveraging collaborative knowledge and addressing interaction sparsity in conversational datasets. We propose an LLM-based framework that uses Reddit-ML32M to align LLM-generated recommendations with CF embeddings, refining rankings for better performance. We evaluate our framework against three sets of baselines: CF-based recommenders using only interactions from CRS tasks, traditional CRS models, and LLM-based methods relying on conversational context without item representations. Our approach achieves consistent improvements, including a 12.32% increase in Hit Rate and a 9.9% improvement in NDCG, outperforming the best-performing baseline that relies on conversational context but lacks collaborative item representations.

对话推荐协同过滤大模型数据关联

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。