让图文模型学会理解商品间的关联关系,提升跨模态匹配效果
SLIP: Structural-aware Language-Image Pretraining for Vision-Language Alignment
- 引入结构对比损失,利用商品共购图建模图文间关系
- 在零样本和少样本任务中均超越CLIP,跨模态检索准确率提升2.1%~4.3%
- 适合做电商、推荐系统等需理解实体关系的视觉语言任务
视觉-语言预训练(VLP)在多个下游任务中取得显著进展,但主要依赖大规模数据。现有方法将图像-文本对视为独立样本,忽略了如电商商品共购图、社交推荐网络中存在的丰富关系结构。受神经科学启发,人类通过关系认知地图编码知识,我们提出结构感知的图文预训练方法SLIP。SLIP引入结构对比损失,在对齐模态的同时建模结构图中邻近实体间的关系。为支持该范式,我们构建了大规模亚马逊商品共购多模态图数据集,实现规模化跨模态结构监督。实验表明,SLIP在零样本和少样本设置下,于跨模态检索与分类任务中持续优于CLIP,验证了关系监督对跨模态对齐的价值。
原文摘要 · Abstract (English)
Vision-Language Pretraining (VLP) has achieved remarkable success across various downstream tasks, but such gains are largely driven by scaling up on training data. Yet, literature methods treat image-text pairs as isolated training examples; this neglects the rich relational structure naturally present in many domains, such as e-commerce product co-purchase graphs and social recommendation networks. Inspired by neuroscientific evidence that human encodes knowledge as relationship cognitive maps, we introduce Structure-aware Language-Image Pretraining (SLIP). SLIP integrates a structural contrastive loss to align modalities while also modeling relationships between neighboring entities in a structured graph. To support this paradigm, we construct a large-scale Amazon Product Co-purchase Multimodal Graph Dataset, enabling structured cross-modality supervision at scale. Experiment results show that SLIP consistently outperforms CLIP on cross-modal retrieval and classification tasks in both zero-shot and few-shot settings, showing the value of relational supervision for cross-modal alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。