发现视觉语言模型在细粒度分类上表现弱,关键在于视觉编码器和预训练策略。
Understanding the Fine-Grained Knowledge Capabilities of Vision-Language Models
- 对比多种视觉语言模型,发现视觉编码器影响细粒度识别性能
- 预训练阶段若解冻语言模型权重,能显著提升细粒度能力
- 改进大语言模型对所有任务帮助均等,但视觉编码器对细粒度更关键
视觉语言模型(VLMs)在视觉问答、文档理解与多模态对话等任务中取得显著进展,涵盖多种基础模型、对齐架构与训练数据。然而,近期研究显示,这些模型在传统图像分类基准上的表现落后,尤其在细粒度视觉知识测试中表现不佳。本文在多个最新VLMs上评估其在细粒度分类任务的表现,揭示了细粒度知识与其它视觉任务之间的差距成因。通过一系列消融实验发现:使用更强的LLM可同等提升所有基准分数;而更优的视觉编码器则显著增强细粒度分类性能。此外,预训练阶段至关重要,特别是当语言模型权重在预训练期间未被冻结时。这些发现为提升VLM的细粒度视觉理解与以视觉为中心的能力提供了新方向。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have made substantial progress across a wide range of visual question answering benchmarks, spanning visual reasoning, document understanding, and multimodal dialogue. These improvements are evident in a wide range of VLMs built on a variety of base models, alignment architectures, and training data. However, recent works show that these models trail behind in traditional image classification benchmarks, which test fine-grained visual knowledge. We test a large number of recent VLMs on fine-grained classification benchmarks and identify potential factors in the disconnect between fine-grained knowledge and other vision benchmarks. Through a series of ablation experiments, we find that using a better LLM improves all benchmark scores equally, while a better vision encoder disproportionately improves fine-grained classification performance. Furthermore, we find that the pretraining stage is also vital to fine-grained performance, particularly when the language model weights are unfrozen during pretraining. These insights pave the way for enhancing fine-grained visual understanding and vision-centric capabilities in VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。