通过结构化绑定提升视觉语言模型对复杂场景的理解能力
Object-centric Binding in Contrastive Language-Image Pretraining
- 引入场景图与槽结构图像表示的绑定模块,实现跨模态结构对齐
- 在多个多物体组合场景数据集上显著提升零样本识别准确率
- 无需硬负样本,适合需要高效理解复杂语义关系的应用
近期视觉语言模型(VLM)的发展主要依赖于对比学习模型如CLIP,其通过将视觉信息与文本描述关联来学习。然而,这类模型在理解包含多个物体及其空间关系的复杂组合场景时存在局限。为此,本文提出一种新方法,摒弃传统依赖硬负样本增强的设计策略,转而将归纳偏置融入预训练的CLIP类模型中,以提升其组合理解能力,且不需额外硬负样本。具体地,我们设计了一个绑定模块,将由文本描述生成的场景图与槽结构图像表示相连接,实现两模态间的结构化相似性评估。同时,利用关系作为文本条件化的视觉约束,更有效地捕捉物体间的复杂交互与上下文关系。实验表明,该模型显著提升了基于CLIP的模型在多物体组合理解任务上的表现,并为复杂场景下更精准、样本高效的图像-文本匹配提供了新路径。
原文摘要 · Abstract (English)
Recent advances in vision language models (VLM) have been driven by contrastive models such as CLIP, which learn to associate visual information with their corresponding text descriptions. However, these models have limitations in understanding complex compositional scenes involving multiple objects and their spatial relationships. To address these challenges, we propose a novel approach that diverges from commonly used strategies, which rely on the design of hard-negative augmentations. Instead, our work focuses on integrating inductive biases into pre-trained CLIP-like models to improve their compositional understanding without using any additional hard-negatives. To that end, we introduce a binding module that connects a scene graph, derived from a text description, with a slot-structured image representation, facilitating a structured similarity assessment between the two modalities. We also leverage relationships as text-conditioned visual constraints, thereby capturing the intricate interactions between objects and their contextual relationships more effectively. Our resulting model not only enhances the performance of CLIP-based models in multi-object compositional understanding but also paves the way towards more accurate and sample-efficient image-text matching of complex scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。