通过指令编辑数据和长描述句提升CLIP的细粒度视觉理解能力。
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
- 利用指令编辑数据生成难负样本,增强模型区分细微差异的能力。
- 引入长描述句与旋转位置编码,捕捉更丰富的语义上下文信息。
- 显著降低多模态大模型的视觉幻觉,适合需要精准视觉推理的场景。
尽管视觉语言模型(如CLIP)在对齐视觉与语言方面取得成功,其在细节化、细粒度视觉理解方面的表现仍面临挑战。我们提出CLIP-IN框架,通过两项核心创新增强CLIP的细粒度感知能力。首先,利用原本用于图像操作的指令编辑数据集作为硬负样本的独特来源,结合对称硬负对比损失,使模型能有效区分细微的视觉-语义差异。其次,引入长描述句,并采用旋转位置编码以捕捉标准CLIP常忽略的丰富语义上下文。实验表明,CLIP-IN在MMVP基准及多个细粒度视觉识别任务中实现显著提升,且不损害其在更广泛分类与检索任务上的零样本鲁棒性。关键的是,将CLIP-IN的视觉表征融入多模态大语言模型后,显著减少视觉幻觉并增强推理能力。本工作凸显了结合针对性指令式对比学习与全面描述信息,在提升视觉语言模型细粒度理解方面的巨大潜力。
原文摘要 · Abstract (English)
Despite the success of Vision-Language Models (VLMs) like CLIP in aligning vision and language, their proficiency in detailed, fine-grained visual comprehension remains a key challenge. We present CLIP-IN, a novel framework that bolsters CLIP's fine-grained perception through two core innovations. Firstly, we leverage instruction-editing datasets, originally designed for image manipulation, as a unique source of hard negative image-text pairs. Coupled with a symmetric hard negative contrastive loss, this enables the model to effectively distinguish subtle visual-semantic differences. Secondly, CLIP-IN incorporates long descriptive captions, utilizing rotary positional encodings to capture rich semantic context often missed by standard CLIP. Our experiments demonstrate that CLIP-IN achieves substantial gains on the MMVP benchmark and various fine-grained visual recognition tasks, without compromising robust zero-shot performance on broader classification and retrieval tasks. Critically, integrating CLIP-IN's visual representations into Multimodal Large Language Models significantly reduces visual hallucinations and enhances reasoning abilities. This work underscores the considerable potential of synergizing targeted, instruction-based contrastive learning with comprehensive descriptive information to elevate the fine-grained understanding of VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。