多模态生成推荐突破单一文本限制,性能提升超20%。
Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics
- 构建多模态生成推荐框架,融合文本与视觉等多源信息。
- 引入对比对齐与模态标识符,实现跨模态语义融合,性能提升20%以上。
- 适合关注多模态推荐、生成式系统优化的研究者与工程师。
生成式推荐(GR)作为推荐系统的新范式,通过隐式关联模态与语义来构建物品表征,区别于以往依赖非语义标识符的自回归模型。然而,现有研究大多将模态孤立处理,通常假设物品内容为单模态(如文本)。我们指出,这一假设在真实世界数据中存在显著局限,且生成模型对模态选择极为敏感。本文聚焦多模态生成推荐(MGR),探讨模态选择的重要性及多模态下有效建模的挑战。通过评估多种多模态融合策略,我们识别关键瓶颈,并提出MGR-LF++——一种增强型晚期融合框架,采用对比模态对齐和特殊标记表示不同模态,相比单模态方案性能提升超过20%。
原文摘要 · Abstract (English)
Generative recommendation (GR) has become a powerful paradigm in recommendation systems that implicitly links modality and semantics to item representation, in contrast to previous methods that relied on non-semantic item identifiers in autoregressive models. However, previous research has predominantly treated modalities in isolation, typically assuming item content is unimodal (usually text). We argue that this is a significant limitation given the rich, multimodal nature of real-world data and the potential sensitivity of GR models to modality choices and usage. Our work aims to explore the critical problem of Multimodal Generative Recommendation (MGR), highlighting the importance of modality choices in GR nframeworks. We reveal that GR models are particularly sensitive to different modalities and examine the challenges in achieving effective GR when multiple modalities are available. By evaluating design strategies for effectively leveraging multiple modalities, we identify key challenges and introduce MGR-LF++, an enhanced late fusion framework that employs contrastive modality alignment and special tokens to denote different modalities, achieving a performance improvement of over 20% compared to single-modality alternatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。