建模多模态差异,提升序列推荐准确率
Multimodal Difference Learning for Sequential Recommendation
- 构建模态感知的物品关系图,增强物品表征
- 设计兴趣集中注意力机制,独立建模多模态用户偏好
- 在5个真实数据集上优于主流方法,适合多模态推荐场景
序列推荐通过建模用户历史行为预测下一个物品,近年来受到广泛关注。随着互联网平台上多模态数据(如图像、文本)的爆发式增长,序列推荐也受益于多模态信息的引入。现有方法通常将物品的模态特征作为辅助信息简单拼接,以学习统一的用户兴趣,但难以捕捉模态间的差异。本文认为,用户兴趣与物品间关系在不同模态下存在差异。为此,提出一种新的多模态差异学习框架MDSRec。首先,利用行为信号构建模态感知的物品关系图,以增强物品表征;其次,设计兴趣集中注意力机制,分别建模不同模态下的用户序列表示;最后,融合多模态用户嵌入实现精准推荐。在五个真实数据集上的实验表明,MDSRec显著优于当前最优基线,验证了多模态差异学习的有效性。
原文摘要 · Abstract (English)
Sequential recommendations have drawn significant attention in modeling the user's historical behaviors to predict the next item. With the booming development of multimodal data (e.g., image, text) on internet platforms, sequential recommendation also benefits from the incorporation of multimodal data. Most methods introduce modal features of items as side information and simply concatenates them to learn unified user interests. Nevertheless, these methods encounter the limitation in modeling multimodal differences. We argue that user interests and item relationships vary across different modalities. To address this problem, we propose a novel Multimodal Difference Learning framework for Sequential Recommendation, MDSRec for brevity. Specifically, we first explore the differences in item relationships by constructing modal-aware item relation graphs with behavior signal to enhance item representations. Then, to capture the differences in user interests across modalities, we design a interest-centralized attention mechanism to independently model user sequence representations in different modalities. Finally, we fuse the user embeddings from multiple modalities to achieve accurate item recommendation. Experimental results on five real-world datasets demonstrate the superiority of MDSRec over state-of-the-art baselines and the efficacy of multimodal difference learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。