用跨注意力与解耦学习,提升多模态推荐系统精度
CADMR: Cross-Attention and Disentangled Learning for Multimodal Recommender Systems
- 通过解耦学习分离模态特征并保留关联性
- 多头交叉注意力增强用户-物品交互表示
- 在三个数据集上显著优于现有方法
推荐系统中日益丰富的多模态数据为提升推荐准确率和用户满意度提供了新途径。然而,这些系统需应对高维稀疏的用户-物品评分矩阵,仅基于每个用户的少量偏好项目进行矩阵重建存在重大挑战。为此,我们提出CADMR——一种基于自编码器的新型多模态推荐框架。CADMR利用多头交叉注意力机制与解耦学习,有效整合异构多模态数据以重构评分矩阵。首先,通过解耦学习分离模态特异性特征并保持其相互依赖性,从而学习联合潜在表示;随后,应用多头交叉注意力机制,基于学习到的多模态物品潜在表示增强用户-物品交互表示。我们在三个基准数据集上评估了CADMR,结果表明其性能显著优于当前最优方法。
原文摘要 · Abstract (English)
The increasing availability and diversity of multimodal data in recommender systems offer new avenues for enhancing recommendation accuracy and user satisfaction. However, these systems must contend with high-dimensional, sparse user-item rating matrices, where reconstructing the matrix with only small subsets of preferred items for each user poses a significant challenge. To address this, we propose CADMR, a novel autoencoder-based multimodal recommender system framework. CADMR leverages multi-head cross-attention mechanisms and Disentangled Learning to effectively integrate and utilize heterogeneous multimodal data in reconstructing the rating matrix. Our approach first disentangles modality-specific features while preserving their interdependence, thereby learning a joint latent representation. The multi-head cross-attention mechanism is then applied to enhance user-item interaction representations with respect to the learned multimodal item latent representations. We evaluate CADMR on three benchmark datasets, demonstrating significant performance improvements over state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。