通过动态加权多模态信号,让推荐系统更懂用户在不同场景下的偏好。
CAMMSR: Category-Guided Attentive Mixture of Experts for Multimodal Sequential Recommendation
- 用类别引导的专家混合模型,按需分配图文等模态权重。
- 在4个公开数据集上显著优于现有方法,提升推荐准确率。
- 适合研究多模态推荐、个性化系统的人参考。
信息密集环境中多媒体数据爆炸式增长,推动了个性化内容发现的需求,使推荐系统成为重要的被动数据管理方式。多模态序列推荐通过融合文本、图像等多样化物品信息,有效丰富物品表征并深化对用户兴趣的理解。然而,现有模型多依赖启发式融合策略,难以捕捉用户-模态交互的动态性与上下文敏感性。真实场景中,用户对模态的偏好不仅个体差异显著,同一用户在不同物品或类别下也会变化。此外,模态间的协同效应——即联合信号引发单独模态无法触发的兴趣——仍被忽视。为此,我们提出CAMMSR:一种类别引导的注意力混合专家模型,用于多模态序列推荐。其核心是类别引导的注意力混合专家(CAMoE)模块,从多视角学习专用物品表征,并显式建模跨模态协同。该模块通过辅助类别预测任务动态分配模态权重,实现多模态信号的自适应融合。同时,设计模态交换对比学习任务,通过序列级增强提升跨模态表示对齐。在四个公开数据集上的大量实验表明,CAMMSR持续优于最先进基线,验证了其在实现自适应、协同性与用户中心化多模态序列推荐中的有效性。
原文摘要 · Abstract (English)
The explosion of multimedia data in information-rich environments has intensified the challenges of personalized content discovery, positioning recommendation systems as an essential form of passive data management. Multimodal sequential recommendation, which leverages diverse item information such as text and images, has shown great promise in enriching item representations and deepening the understanding of user interests. However, most existing models rely on heuristic fusion strategies that fail to capture the dynamic and context-sensitive nature of user-modal interactions. In real-world scenarios, user preferences for modalities vary not only across individuals but also within the same user across different items or categories. Moreover, the synergistic effects between modalities-where combined signals trigger user interest in ways isolated modalities cannot-remain largely underexplored. To this end, we propose CAMMSR, a Category-guided Attentive Mixture of Experts model for Multimodal Sequential Recommendation. At its core, CAMMSR introduces a category-guided attentive mixture of experts (CAMoE) module, which learns specialized item representations from multiple perspectives and explicitly models inter-modal synergies. This component dynamically allocates modality weights guided by an auxiliary category prediction task, enabling adaptive fusion of multimodal signals. Additionally, we design a modality swap contrastive learning task to enhance cross-modal representation alignment through sequence-level augmentation. Extensive experiments on four public datasets demonstrate that CAMMSR consistently outperforms state-of-the-art baselines, validating its effectiveness in achieving adaptive, synergistic, and user-centric multimodal sequential recommendation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。