arXiv:2501.14269cs.IRcs.AI2025-01中稿 · WWW 2025被引 48

提出分层时序专家混合模型,提升多模态序列推荐效果

Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential Recommendation

  • 双级专家混合架构分离兴趣相关与时间动态信息
  • 在四个数据集上优于当前最优方法,显著缓解数据稀疏问题
  • 适合需要精准时序建模的电商/视频推荐场景

多模态序列推荐通过融合多源信息学习更全面的用户偏好,已成为学术与工业界关键课题。现有方法虽通过自适应模态融合捕捉用户偏好演变,但常忽略多模态数据中冗余无关信息的干扰,且仅依赖时间顺序的隐式时序信号,未能有效建模动态兴趣。为此,本文提出分层时序感知专家混合模型(HM4SR),采用两级专家混合结构与多任务学习策略:第一级交互专家混合(Interactive MoE)从每项的多模态数据中提取用户兴趣相关特征;第二级时间专家混合(Temporal MoE)引入显式时间嵌入,基于时间戳编码建模用户动态兴趣。为缓解数据稀疏性,设计三项辅助任务:序列级类别预测(CP)增强物品特征理解,基于ID的对比学习(IDCL)对齐序列上下文与用户兴趣,以及占位符对比学习(PCL)融合时序与模态信息以建模动态兴趣。在四个公开数据集上的大量实验验证了该方法的有效性,性能优于多个前沿模型。

原文摘要 · Abstract (English)

Multi-modal sequential recommendation (SR) leverages multi-modal data to learn more comprehensive item features and user preferences than traditional SR methods, which has become a critical topic in both academia and industry. Existing methods typically focus on enhancing multi-modal information utility through adaptive modality fusion to capture the evolving of user preference from user-item interaction sequences. However, most of them overlook the interference caused by redundant interest-irrelevant information contained in rich multi-modal data. Additionally, they primarily rely on implicit temporal information based solely on chronological ordering, neglecting explicit temporal signals that could more effectively represent dynamic user interest over time. To address these limitations, we propose a Hierarchical time-aware Mixture of experts for multi-modal Sequential Recommendation (HM4SR) with a two-level Mixture of Experts (MoE) and a multi-task learning strategy. Specifically, the first MoE, named Interactive MoE, extracts essential user interest-related information from the multi-modal data of each item. Then, the second MoE, termed Temporal MoE, captures user dynamic interests by introducing explicit temporal embeddings from timestamps in modality encoding. To further address data sparsity, we propose three auxiliary supervision tasks: sequence-level category prediction (CP) for item feature understanding, contrastive learning on ID (IDCL) to align sequence context with user interests, and placeholder contrastive learning (PCL) to integrate temporal information with modalities for dynamic interest modeling. Extensive experiments on four public datasets verify the effectiveness of HM4SR compared to several state-of-the-art approaches.

序列推荐多模态时序建模MoE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。