发现视觉语言模型看不懂图片变换,影响图像编辑效果。
On the Limitations of Vision-Language Models in Understanding Image Transforms
- 用增强版Flickr8k数据集测试模型对图像变换的理解能力
- CLIP和SigLIP在多种图像变换上表现不佳,准确率低于随机猜测
- 研究结果对图像生成、编辑等应用有重要警示意义
视觉语言模型(VLMs)在图像生成、视觉问答、多模态对话等任务中表现出色,但在基本图像变换理解上存在明显缺陷。本文聚焦于CLIP与SigLIP两个代表性模型的图像级理解能力,发现其对多种图像变换缺乏有效认知。为此,我们构建了增强版Flickr8k数据集,为每张图像添加详细的变换描述。进一步分析表明,该缺陷严重影响下游任务,尤其在图像编辑场景下,当前最先进的Image2Image模型在简单变换任务上的性能显著下降。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) have demonstrated significant potential in various downstream tasks, including Image/Video Generation, Visual Question Answering, Multimodal Chatbots, and Video Understanding. However, these models often struggle with basic image transformations. This paper investigates the image-level understanding of VLMs, specifically CLIP by OpenAI and SigLIP by Google. Our findings reveal that these models lack comprehension of multiple image-level augmentations. To facilitate this study, we created an augmented version of the Flickr8k dataset, pairing each image with a detailed description of the applied transformation. We further explore how this deficiency impacts downstream tasks, particularly in image editing, and evaluate the performance of state-of-the-art Image2Image models on simple transformations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。