通过增强语义图与对比学习提升多模态推荐效果
Semantic Item Graph Enhancement for Multimodal Recommendation
- 融合用户行为信号优化各模态的物品语义图
- 利用模态引导的扰动生成对比视图,增强抗噪能力
- 双阶段对齐机制保证多源表示一致性,适合推荐系统研究者
多模态推荐系统通过利用物品的多模态信息提升了性能。现有方法通常从原始模态特征构建特定模态的物品-物品语义图,并与用户-物品交互图协同使用以增强用户偏好学习。然而,这些语义图存在语义不足问题,包括(1)物品间协同信号建模不充分;(2)原始模态特征中的噪声引入结构失真,最终影响性能。为此,我们首先从交互图中提取协同信号,并注入到各模态特定的物品语义图中以增强语义建模。接着,设计一种基于模态的个性化嵌入扰动机制,通过模态引导的个性化强度注入扰动,生成对比视图,使模型能通过对比学习获得抗噪表示,从而降低语义图中结构噪声的影响。此外,提出双表示对齐机制:首先使用行为表示作为锚点,通过锚定式InfoNCE损失对齐多个语义表示;再通过标准InfoNCE将行为表示与融合语义对齐,确保表示一致性。在四个基准数据集上的大量实验验证了所提框架的有效性。
原文摘要 · Abstract (English)
Multimodal recommendation systems have attracted increasing attention for their improved performance by leveraging items' multimodal information. Prior methods often build modality-specific item-item semantic graphs from raw modality features and use them as supplementary structures alongside the user-item interaction graph to enhance user preference learning. However, these semantic graphs suffer from semantic deficiencies, including (1) insufficient modeling of collaborative signals among items and (2) structural distortions introduced by noise in raw modality features, ultimately compromising performance. To address these issues, we first extract collaborative signals from the interaction graph and infuse them into each modality-specific item semantic graph to enhance semantic modeling. Then, we design a modulus-based personalized embedding perturbation mechanism that injects perturbations with modulus-guided personalized intensity into embeddings to generate contrastive views. This enables the model to learn noise-robust representations through contrastive learning, thereby reducing the effect of structural noise in semantic graphs. Besides, we propose a dual representation alignment mechanism that first aligns multiple semantic representations via a designed Anchor-based InfoNCE loss using behavior representations as anchors, and then aligns behavior representations with the fused semantics by standard InfoNCE, to ensure representation consistency. Extensive experiments on four benchmark datasets validate the effectiveness of our framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。