arXiv:2508.00452cs.IRcs.AI2025-08AAAI被引 2

用多模态多视角生成模型解决新商品推荐难题

M^2VAE: Multi-Modal Multi-View Variational Autoencoder for Cold-start Item Recommendation

  • 为物品属性、图像等设计类型特异性隐变量,融合共性与个性特征
  • 在多个真实数据集上显著优于基线方法,尤其在冷启动场景提升明显
  • 适合做推荐系统研究或工业应用中冷启动问题的开发者

冷启动物品推荐是推荐系统中的重大挑战,尤其当新物品缺乏历史交互数据时。现有方法虽利用多模态内容缓解该问题,但常忽略模态内在的多视角结构以及共享与模态特有特征的区别。本文提出多模态多视角变分自编码器(M^2VAE),一种生成模型,用于建模属性和多模态特征中的共性与独特视图,以及用户对单一类型物品特征的偏好。具体而言,我们为物品ID、类别属性和图像特征生成类型特异性隐变量,并使用专家乘积(PoE)推导公共表示。通过解耦对比损失将共性视图与独特视图分离,同时保持特征信息量。为建模用户倾向,采用偏好引导的专家混合(MoE)自适应融合表示。进一步通过对比学习引入共现信号,无需预训练。在多个真实数据集上的实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Cold-start item recommendation is a significant challenge in recommendation systems, particularly when new items are introduced without any historical interaction data. While existing methods leverage multi-modal content to alleviate the cold-start issue, they often neglect the inherent multi-view structure of modalities, the distinction between shared and modality-specific features. In this paper, we propose Multi-Modal Multi-View Variational AutoEncoder (M^2VAE), a generative model that addresses the challenges of modeling common and unique views in attribute and multi-modal features, as well as user preferences over single-typed item features. Specifically, we generate type-specific latent variables for item IDs, categorical attributes, and image features, and use Product-of-Experts (PoE) to derive a common representation. A disentangled contrastive loss decouples the common view from unique views while preserving feature informativeness. To model user inclinations, we employ a preference-guided Mixture-of-Experts (MoE) to adaptively fuse representations. We further incorporate co-occurrence signals via contrastive learning, eliminating the need for pretraining. Extensive experiments on real-world datasets validate the effectiveness of our approach.

推荐系统冷启动多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。