视觉信息有用吗?看懂上下文再决定。
Deciding When to Rely on Visual Information: Gated Multimodal Fusion in Sequential Recommendation

- 根据物品和用户历史动态决定是否用视觉信息
- 在用户互动少时,视觉信息作用更大
- 能解释哪些物品适合用视觉推荐
多模态序列推荐系统通常统一融合视觉与协同信号,将视觉特征视为通用信息。我们提出,视觉效用(即视觉信号对推荐质量的贡献)是依赖物品和用户交互历史的潜在变量,而非固定属性。为此,我们提出VisGate框架,基于物品嵌入和用户当前序列上下文,做出自适应的物品级融合决策。视觉表示通过序列共现模式的对比学习获得,保持与协同嵌入的互补性,而非对齐到共享空间。实验表明,该方法性能媲美先进模型;进一步分析发现:视觉效用因物品而异,在协同信号稀疏时提升,且与视觉独特性在语义上有意义的相关性。这些结果强调细粒度融合与模态互补的重要性,并证明可通过学习的门控行为估算和解释物品级视觉效用。
原文摘要 · Abstract (English)
Multimodal sequential recommender systems commonly fuse visual and collaborative signals uniformly, treating visual features as generically informative regardless of item or user context. We argue that visual utility, defined as the contribution of visual signals to recommendation quality, is a latent contextual variable that depends on both the item and the user's interaction history rather than a fixed item property. To model this variability, we introduce VisGate, a framework that makes adaptive item-level fusion decisions conditioned on item embeddings and the user's current sequence context. Visual representations are learned through a contrastive objective over sequential co-occurrence patterns, preserving complementarity with collaborative embeddings rather than aligning them into a shared space. Beyond achieving competitive recommendation performance, VisGate's learned gate serves as a measurement tool for understanding when and why visual information is beneficial. Our analyses show that visual utility varies across items, increases under interaction sparsity when collaborative signals are weak, and correlates with visual distinctiveness in semantically meaningful ways. Together, these findings highlight the importance of both fine-grained fusion and modality complementarity, while demonstrating that item-level visual utility can be estimated and interpreted through learned gating behaviour.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。