评测大模型多层级视觉感知能力,发现其高阶理解远不如人类。
MVP-Bench: Can Large Vision--Language Models Conduct Multi-level Visual Perception Like Humans?
- 构建MVP-Bench基准,涵盖自然与合成图像的多层级视觉感知测试
- GPT-4o在高阶任务准确率仅56%,低于低阶任务的74%
- 模型对合成图像语义理解能力差,难以类人泛化
人类在多个层次上进行视觉感知,包括低层次的对象识别和高层次的语义理解(如行为解读)。低层次细节的细微差异可能导致高层次感知的显著变化。例如,将人物手中的购物袋替换为枪支,会暗示暴力行为,进而关联到犯罪或攻击性活动。尽管多模态任务取得进展,但大型视觉语言模型(LVLMs)在多层级视觉感知方面仍缺乏系统研究。为此,我们提出MVP-Bench,首个系统评估LVLM在低、高阶视觉感知能力的基准。该基准涵盖自然与合成图像,探究内容修改对模型感知的影响。通过MVP-Bench,我们诊断了10个开源和2个闭源LVLM的感知表现,结果显示高阶感知任务严重挑战现有模型。最先进模型GPT-4o在是非题上准确率仅为56%,而低阶场景下可达74%。此外,自然图像与篡改图像间的性能差距表明,当前模型无法像人类一样泛化理解合成图像的视觉语义。数据与代码已公开于https://github.com/GuanzhenLi/MVP-Bench。
原文摘要 · Abstract (English)
Humans perform visual perception at multiple levels, including low-level object recognition and high-level semantic interpretation such as behavior understanding. Subtle differences in low-level details can lead to substantial changes in high-level perception. For example, substituting the shopping bag held by a person with a gun suggests violent behavior, implying criminal or violent activity. Despite significant advancements in various multimodal tasks, Large Visual-Language Models (LVLMs) remain unexplored in their capabilities to conduct such multi-level visual perceptions. To investigate the perception gap between LVLMs and humans, we introduce MVP-Bench, the first visual-language benchmark systematically evaluating both low- and high-level visual perception of LVLMs. We construct MVP-Bench across natural and synthetic images to investigate how manipulated content influences model perception. Using MVP-Bench, we diagnose the visual perception of 10 open-source and 2 closed-source LVLMs, showing that high-level perception tasks significantly challenge existing LVLMs. The state-of-the-art GPT-4o only achieves an accuracy of $56\%$ on Yes/No questions, compared with $74\%$ in low-level scenarios. Furthermore, the performance gap between natural and manipulated images indicates that current LVLMs do not generalize in understanding the visual semantics of synthetic images as humans do. Our data and code are publicly available at https://github.com/GuanzhenLi/MVP-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。