构建基础视觉技能数据集,揭示大模型在简单几何任务上的短板。
Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
- 将2D几何视觉理解拆解为不可再分的原子技能
- 在新数据集上测试顶尖模型,准确率远低于人类水平
- 适合关注视觉认知本质与模型评测的科研人员
近期视觉语言模型(VLMs)在多模态理解与推理方面表现优异,但在看似简单的视觉任务上仍表现不佳。本文聚焦于基础二维欧几里得几何领域,系统分类了基本且不可分割的视觉感知能力,称为原子视觉技能。为此,我们构建了原子视觉技能数据集(AVSD),用于评估VLM在这些原子技能上的表现。基于AVSD,我们对当前最先进的VLM进行了基准测试,发现它们在这些任务上表现糟糕,尽管对成人而言这些任务极为简单。研究结果凸显了针对原子级而非复合型视觉感知任务设计专用数据集的必要性。
原文摘要 · Abstract (English)
Recent Vision-Language Models (VLMs) have demonstrated impressive multimodal comprehension and reasoning capabilities, yet they often struggle with trivially simple visual tasks. In this work, we focus on the domain of basic 2D Euclidean geometry and systematically categorize the fundamental, indivisible visual perception skills, which we refer to as atomic visual skills. We then introduce the Atomic Visual Skills Dataset (AVSD) for evaluating VLMs on the atomic visual skills. Using AVSD, we benchmark state-of-the-art VLMs and find that they struggle with these tasks, despite being trivial for adult humans. Our findings highlight the need for purpose-built datasets to train and evaluate VLMs on atomic, rather than composite, visual perception tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。