用视觉错觉测试模型,发现加模糊滤波能显著提升识别能力。
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
- 构建四类错觉数据集,评估多模态模型对幻觉图像的理解能力。
- 无微调下BLIP-2在错觉动物数据集上超越人类表现。
- 用高斯与模糊低通滤波器增强模型鲁棒性,方法简单有效。
近年来,视觉问答(VQA)取得显著进展,尤其得益于融合视觉与语言理解的多模态模型。然而,现有VQA数据集常忽略图像错觉带来的复杂挑战,这对人类感知和模型解读均构成独特考验。本文提出新的任务——错觉视觉问答(Illusory VQA),并构建四个专用数据集:IllusionMNIST、IllusionFashionMNIST、IllusionAnimals 和 IllusionChar,用于评估先进多模态模型在识别与解释视觉错觉方面的能力。我们评估了多种模型的零样本性能,对部分模型进行微调,并提出一种基于高斯与模糊低通滤波器的简单有效错觉检测方法。实验表明,该方法显著提升模型性能;在未微调情况下,BLIP-2 在 IllusionAnimals 上的表现甚至超过人类。研究揭示了人类与模型在错觉感知上的差异,并证明微调与特定预处理技术可大幅增强模型鲁棒性。本工作推动多模态模型向更类人视觉理解发展,提示未来可探索可学习参数的滤波器优化方向。
原文摘要 · Abstract (English)
In recent years, Visual Question Answering (VQA) has made significant strides, particularly with the advent of multimodal models that integrate vision and language understanding. However, existing VQA datasets often overlook the complexities introduced by image illusions, which pose unique challenges for both human perception and model interpretation. In this study, we introduce a novel task called Illusory VQA, along with four specialized datasets: IllusionMNIST, IllusionFashionMNIST, IllusionAnimals, and IllusionChar. These datasets are designed to evaluate the performance of state-of-the-art multimodal models in recognizing and interpreting visual illusions. We assess the zero-shot performance of various models, fine-tune selected models on our datasets, and propose a simple yet effective solution for illusion detection using Gaussian and blur low-pass filters. We show that this method increases the performance of models significantly and in the case of BLIP-2 on IllusionAnimals without any fine-tuning, it outperforms humans. Our findings highlight the disparity between human and model perception of illusions and demonstrate that fine-tuning and specific preprocessing techniques can significantly enhance model robustness. This work contributes to the development of more human-like visual understanding in multimodal models and suggests future directions for adapting filters using learnable parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。