通过行为层面的跨模态兴趣融合,解决推荐系统中物品ID稀疏问题。
Distribution-Guided Auto-Encoder for User Multimodal Interest Cross Fusion
- 基于用户行为序列动态引导多模态特征融合
- 在MovieLens和Amazon-Book数据集上显著提升推荐效果
- 适合研究多模态推荐与用户行为建模的读者
传统推荐方法依赖物品ID嵌入向量捕捉隐式协同过滤信号,但因ID特征稀疏性导致数据稀疏问题。为缓解此问题,现有模型引入多模态物品信息以提升推荐精度。然而,现有方法多采用早期融合策略,仅关注文本与图像特征的组合,忽视用户行为序列的上下文影响,难以根据行为模式动态调整多模态兴趣表示,限制了对用户多模态兴趣的精准建模。为此,本文提出分布引导的多模态兴趣自编码器(Distribution-Guided Multimodal-Interest Auto-Encoder, DMAE),实现行为层面的用户多模态兴趣跨融合。大量实验表明,DMAE在MovieLens和Amazon-Book数据集上均显著优于基线模型。
原文摘要 · Abstract (English)
Traditional recommendation methods rely on correlating the embedding vectors of item IDs to capture implicit collaborative filtering signals to model the user's interest in the target item. Consequently, traditional ID-based methods often encounter data sparsity problems stemming from the sparse nature of ID features. To alleviate the problem of item ID sparsity, recommendation models incorporate multimodal item information to enhance recommendation accuracy. However, existing multimodal recommendation methods typically employ early fusion approaches, which focus primarily on combining text and image features, while neglecting the contextual influence of user behavior sequences. This oversight prevents dynamic adaptation of multimodal interest representations based on behavioral patterns, consequently restricting the model's capacity to effectively capture user multimodal interests. Therefore, this paper proposes the Distribution-Guided Multimodal-Interest Auto-Encoder (DMAE), which achieves the cross fusion of user multimodal interest at the behavioral level.Ultimately, extensive experiments demonstrate the superiority of DMAE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。