arXiv:2512.21543cs.IR2025-12被引 2

用协同信号增强多模态融合,提升推荐生成效果

CEMG: Collaborative-Enhanced Multimodal Generative Recommendation

  • 用协同信号引导视觉文本特征动态融合
  • 将融合结果转为离散语义码,提升表示质量
  • 适合需要高质量多模态推荐的场景

生成式推荐模型常面临两大挑战:(1) 协同信号整合浅层化,(2) 多模态特征解耦融合。这些限制阻碍了完整物品表示的构建。为此,我们提出 CEMG 框架,包含三个阶段:首先通过多模态融合层,在协同信号引导下动态整合视觉与文本特征;其次通过统一模态标记化阶段,采用残差量化变分自编码器(RQ-VAE)将融合表示转换为离散语义码;最后在端到端生成推荐阶段,微调大语言模型以自回归方式生成这些物品代码。大量实验表明,CEMG 显著优于现有最先进基线方法。

原文摘要 · Abstract (English)

Generative recommendation models often struggle with two key challenges: (1) the superficial integration of collaborative signals, and (2) the decoupled fusion of multimodal features. These limitations hinder the creation of a truly holistic item representation. To overcome this, we propose CEMG, a novel Collaborative-Enhaned Multimodal Generative Recommendation framework. Our approach features a Multimodal Fusion Layer that dynamically integrates visual and textual features under the guidance of collaborative signals. Subsequently, a Unified Modality Tokenization stage employs a Residual Quantization VAE (RQ-VAE) to convert this fused representation into discrete semantic codes. Finally, in the End-to-End Generative Recommendation stage, a large language model is fine-tuned to autoregressively generate these item codes. Extensive experiments demonstrate that CEMG significantly outperforms state-of-the-art baselines.

推荐系统多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。