arXiv:2511.02113cs.IR2025-11被引 1

用视觉语言模型和信息分解提升多模态推荐效果

Enhancing Multimodal Recommendations with Vision-Language Models and Information-Aware Fusion

  • 用标题引导生成细粒度图像描述,实现视觉与文本对齐
  • 基于信息分解的融合机制,显著提升视觉特征贡献度
  • 在三个亚马逊数据集上超越主流方法,适合多模态推荐研究者

近年来,多模态推荐(MMR)通过整合视觉与文本内容来丰富物品表征。然而,现有方法常依赖粗粒度视觉特征和简单融合策略,导致表征冗余或错位。从信息论角度,有效融合需平衡各模态的独有、共享与冗余信息,以保留互补线索。为此,我们提出VIRAL框架,包含两个组件:(i) 基于视觉语言模型(VLM)的视觉增强模块,生成与标题引导的细粒度描述,实现语义对齐的图像表征;(ii) 受部分信息分解(PID)启发的信息感知融合模块,用于解耦并整合互补信号。在三个Amazon数据集上的实验表明,VIRAL持续优于强基线方法,且显著提升视觉特征的贡献度。

原文摘要 · Abstract (English)

Recent advances in multimodal recommendation (MMR) highlight the potential of integrating visual and textual content to enrich item representations. However, existing methods often rely on coarse visual features and naive fusion strategies, resulting in redundant or misaligned representations. From an information-theoretic perspective, effective fusion should balance unique, shared, and redundant modality information to preserve complementary cues. To this end, we propose VIRAL, a novel Vision-Language and Information-aware Recommendation framework that enhances multimodal fusion through two components: (i) a VLM-based visual enrichment module that generates fine-grained, title-guided descriptions for semantically aligned image representations, and (ii) an information-aware fusion module inspired by Partial Information Decomposition (PID) to disentangle and integrate complementary signals. Experiments on three Amazon datasets show that VIRAL consistently outperforms strong multimodal baselines and substantially improves the contribution of visual features.

多模态推荐视觉语言模型信息融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。