融合多模态信息与稀疏注意力,提升长序列推荐精度
Multimodal Fusion And Sparse Attention-based Alignment Model for Long Sequential Recommendation
- 用标题作为跨模态语义锚,联合优化四类损失实现高质量融合
- 多粒度稀疏注意力机制支持长序列建模,捕捉用户兴趣演化
- 适合需要精准理解用户长期偏好的推荐系统场景
多模态推荐虽能增强物品理解,但如何有效利用多模态行为序列并挖掘用户多尺度兴趣,仍面临挑战。为此,我们提出MUFASA模型,包含两个核心组件:首先,多模态融合层(MFL)以物品标题为跨类型语义锚点,通过四项定制损失联合训练,促进跨类型语义对齐、协同空间对齐、标题相似性结构保留及融合空间分布正则化,生成高质量融合表示;其次,稀疏注意力引导对齐层(SAL)采用窗口注意力、块级注意力与选择性注意力的多粒度稀疏机制,分层建模用户兴趣的演化与块内细粒度变化,输出鲁棒的用户与物品表示。在真实数据集上的实验表明,MUFASA持续超越现有基线,线上A/B测试也验证了其在生产环境中的显著效果,证明其在利用多模态信号和准确捕捉多样化用户偏好方面的有效性。
原文摘要 · Abstract (English)
Recent advances in multimodal recommendation enable richer item understanding, while modeling users' multi-scale interests across temporal horizons has attracted growing attention. However, effectively exploiting multimodal item sequences and mining multi-grained user interests to substantially bridge the gap between content comprehension and recommendation remain challenging. To address these issues, we propose MUFASA, a MUltimodal Fusion And Sparse Attention-based Alignment model for long sequential recommendation. Our model comprises two core components. First, the Multimodal Fusion Layer (MFL) leverages item titles as a cross-genre semantic anchor and is trained with a joint objective of four tailored losses that promote: (i) cross-genre semantic alignment, (ii) alignment to the collaborative space for recommendation, (iii) preserving the similarity structure defined by titles and preventing modality representation collapse, and (iv) distributional regularization of the fusion space. This yields high-quality fused item representations for further preference alignment. Second, the Sparse Attention-guided Alignment Layer (SAL) scales to long user-behavior sequences via a multi-granularity sparse attention mechanism, which incorporates windowed attention, block-level attention, and selective attention, to capture user interests hierarchically and across temporal horizons. SAL explicitly models both the evolution of coherent interest blocks and fine-grained intra-block variations, producing robust user and item representations. Extensive experiments on real-world benchmarks show that MUFASA consistently surpasses state-of-the-art baselines. Moreover, online A/B tests demonstrate significant gains in production, confirming MUFASA's effectiveness in leveraging multimodal cues and accurately capturing diverse user preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。