用核方法对齐CLIP与DINOv2,提升视觉模型细粒度感知能力。
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
- 基于核函数设计新对齐目标,实现图像编码器间高效优化
- 零样本物体识别与空间定位任务性能显著提升
- 适用于需增强视觉理解的多模态大模型下游应用
视觉语言模型(如CLIP)在对齐视觉与文本表征方面取得显著进展,已成为许多多模态大语言模型(MLLMs)如LLaVA和OpenFlamingo的核心组件。然而,多项研究指出,CLIP在细粒度感知方面存在局限,导致下游MLLMs表现严重受限。相比之下,以视觉为中心的基础模型(如DINOv2)在捕捉图像细节方面表现出色。本文提出一种新型核基对齐方法,将CLIP的视觉表征与DINOv2对齐,在保持与文本嵌入兼容性的同时增强感知能力。该对齐目标支持高效的随机优化。在仅图像端微调后,视觉编码器保持与冻结文本编码器的兼容性,并在零样本物体识别、细粒度空间推理与定位任务中取得显著提升。通过集成该对齐后的视觉编码器,下游MLLMs性能也得到增强。
原文摘要 · Abstract (English)
Vision-language models, such as CLIP, have achieved significant success in aligning visual and textual representations, becoming essential components of many multi-modal large language models (MLLMs) like LLaVA and OpenFlamingo. However, numerous studies have identified CLIP's limited fine-grained perception as a critical drawback, leading to substantial failures in downstream MLLMs. In contrast, vision-centric foundation models like DINOv2 demonstrate remarkable capabilities in capturing fine details from images. In this work, we propose a novel kernel-based method to align CLIP's visual representation with that of DINOv2, ensuring that the resulting embeddings maintain compatibility with text embeddings while enhancing perceptual capabilities. Our alignment objective is designed for efficient stochastic optimization. Following this image-only alignment fine-tuning, the visual encoder retains compatibility with the frozen text encoder and exhibits significant improvements in zero-shot object recognition, fine-grained spatial reasoning, and localization. By integrating the aligned visual encoder, downstream MLLMs also demonstrate enhanced performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。