用多模态信息提升生成式推荐的语义标识质量
Multi-Aspect Cross-modal Quantization for Generative Recommendation
- 跨模态量化构建低冲突语义标识
- 多角度对齐提升生成模型效果,准确率显著提高
- 适合需要高精度推荐的多模态场景
生成式推荐(GR)作为推荐系统的新范式,依赖量化表示将物品特征离散化,将用户历史交互建模为离散标记序列,并通过下一标记预测方法预测下一个物品。其挑战在于构建层次化、低冲突且有利于模型训练的高质量语义标识(IDs)。现有方法难以有效利用多模态信息,也未能捕捉不同模态间的深层复杂交互,制约了高质量语义标识的学习与模型训练。为此,我们提出多方面跨模态量化生成推荐框架(MACRec),在语义标识学习和生成模型训练中引入多模态信息。首先,在标识学习阶段引入跨模态量化,通过多模态互补融合降低冲突率,提升码本可用性;其次,引入多方面跨模态对齐机制,包括隐式与显式对齐,进一步增强生成能力。我们在三个知名推荐数据集上进行了大量实验,验证了所提方法的有效性。
原文摘要 · Abstract (English)
Generative Recommendation (GR) has emerged as a new paradigm in recommender systems. This approach relies on quantized representations to discretize item features, modeling users' historical interactions as sequences of discrete tokens. Based on these tokenized sequences, GR predicts the next item by employing next-token prediction methods. The challenges of GR lie in constructing high-quality semantic identifiers (IDs) that are hierarchically organized, minimally conflicting, and conducive to effective generative model training. However, current approaches remain limited in their ability to harness multimodal information and to capture the deep and intricate interactions among diverse modalities, both of which are essential for learning high-quality semantic IDs and for effectively training GR models. To address this, we propose Multi-Aspect Cross-modal quantization for generative Recommendation (MACRec), which introduces multimodal information and incorporates it into both semantic ID learning and generative model training from different aspects. Specifically, we first introduce cross-modal quantization during the ID learning process, which effectively reduces conflict rates and thus improves codebook usability through the complementary integration of multimodal information. In addition, to further enhance the generative ability of our GR model, we incorporate multi-aspect cross-modal alignments, including the implicit and explicit alignments. Finally, we conduct extensive experiments on three well-known recommendation datasets to demonstrate the effectiveness of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。