arXiv:2504.10921cs.IR2025-04被引 16

用多模态图谱增强对话推荐,提升个性化与自然度

MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender Systems

  • 构建文本、图像多模态图结构,捕捉用户偏好深层语义
  • 结合提示学习,显著提升推荐准确率与回复自然性
  • 适合研究对话推荐、多模态融合的开发者参考

对话推荐系统(CRS)通过对话交互提供个性化推荐。现有方法主要依赖对话上下文提取用户偏好,但因对话内容短且稀疏,难以全面捕捉偏好。本文提出多模态语义图提示学习框架MSCRS,首先提取对话中提及物品的文本和图像特征;其次,通过构建模态特定的图结构,捕捉协同、文本、图像等不同模态间的高阶语义关联;最后,创新性地将多模态语义图与提示学习结合,利用大语言模型探索高维语义关系。实验表明,该方法显著提升了项目推荐准确率,并生成更自然、符合上下文的回复内容。

原文摘要 · Abstract (English)

Conversational Recommender Systems (CRSs) aim to provide personalized recommendations by interacting with users through conversations. Most existing studies of CRS focus on extracting user preferences from conversational contexts. However, due to the short and sparse nature of conversational contexts, it is difficult to fully capture user preferences by conversational contexts only. We argue that multi-modal semantic information can enrich user preference expressions from diverse dimensions (e.g., a user preference for a certain movie may stem from its magnificent visual effects and compelling storyline). In this paper, we propose a multi-modal semantic graph prompt learning framework for CRS, named MSCRS. First, we extract textual and image features of items mentioned in the conversational contexts. Second, we capture higher-order semantic associations within different semantic modalities (collaborative, textual, and image) by constructing modality-specific graph structures. Finally, we propose an innovative integration of multi-modal semantic graphs with prompt learning, harnessing the power of large language models to comprehensively explore high-dimensional semantic relationships. Experimental results demonstrate that our proposed method significantly improves accuracy in item recommendation, as well as generates more natural and contextually relevant content in response generation.

对话推荐多模态图神经网络提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。