通过挖掘跨模态协同效应,提升生成式推荐的语义理解能力
SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative Recommendation

- 设计协同生成框架,显式建模跨模态依赖关系
- 在三个基准数据集上超越现有方法,显著提升推荐效果
- 适合关注多模态融合与生成推荐的科研人员
生成式推荐(GR)将物品推荐建模为基于物品标识符的序列到序列生成任务。近期研究引入多模态信号以提供更丰富的粒度证据。然而,现有方法主要依赖对齐中心的融合策略,未充分挖掘模态间的协同信息。实践中,协同信息对捕捉单个模态无法推断的新兴物品属性至关重要,这些属性蕴含内在语义并引导用户偏好,使模型超越表面特征匹配。为此,我们提出SynGR框架,显式促进生成过程中跨模态依赖的利用。通过抑制对主导模态的过度依赖,SynGR能够捕捉超越共享或模态特定信号的新兴物品语义。在三个基准数据集上的大量实验表明,SynGR实现更优性能。
原文摘要 · Abstract (English)
Generative Recommendation (GR) has emerged as a promising paradigm by formulating item recommendation as a sequence-to-sequence generation task over item identifiers. Recent studies have incorporated multimodal signals to provide richer token-level evidence for generation. However, existing approaches largely rely on alignment-centric fusion and underexplore synergistic information across modalities. In practice, synergistic information plays a critical role in capturing emergent item properties that cannot be inferred from any single modality alone. Such properties encode intrinsic item semantics and guide user preferences, enabling models to move beyond surface-level feature matching. To address this limitation, we propose \textbf{SynGR}, a synergistic generative recommendation framework that explicitly encourages the exploitation of cross-modal dependencies during generation. By constraining overreliance on dominant modalities, SynGR enables the model to capture emergent item semantics beyond shared or modality-specific signals. Extensive experiments across three benchmark datasets demonstrate that SynGR achieves superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。