arXiv:2411.06287cs.CV2024-11

测试视觉语言模型对抽象形状的识别能力,发现其仍远不如人类。

Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models

  • 设计新数据集IllusionBench,用元素排列制造抽象形状挑战模型
  • 人类能轻松识别的形状,当前模型识别准确率显著偏低
  • 适合关注视觉感知鲁棒性与模型可解释性的研究者

尽管形状感知在人类视觉中至关重要,早期神经图像分类器却较少依赖形状信息进行物体识别,而更关注其他(常为虚假)特征。尽管近期研究认为当前大型视觉语言模型(VLMs)对形状的依赖程度更高,但我们发现它们在此方面仍存在严重局限。为量化这些局限,我们提出了IllusionBench,一个数据集,通过场景中视觉元素的排列来呈现抽象形状,挑战当前最先进的VLMs。大规模评估显示,尽管人类标注者可轻易识别这些形状,现有VLMs却难以辨识,揭示了未来构建更鲁棒视觉感知系统的重要方向。完整数据集与代码库见: https://arshiahemmat.github.io/illusionbench/

原文摘要 · Abstract (English)

Despite the importance of shape perception in human vision, early neural image classifiers relied less on shape information for object recognition than other (often spurious) features. While recent research suggests that current large Vision-Language Models (VLMs) exhibit more reliance on shape, we find them to still be seriously limited in this regard. To quantify such limitations, we introduce IllusionBench, a dataset that challenges current cutting-edge VLMs to decipher shape information when the shape is represented by an arrangement of visual elements in a scene. Our extensive evaluations reveal that, while these shapes are easily detectable by human annotators, current VLMs struggle to recognize them, indicating important avenues for future work in developing more robust visual perception systems. The full dataset and codebase are available at: \url{https://arshiahemmat.github.io/illusionbench/}

视觉语言模型形状识别感知评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。