测试视觉模型在密集场景下的理解能力,发现当前模型表现远未达标。
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
- 构建高密度画面的VQA数据集,聚焦细节感知与推理。
- 最佳模型仅19.6%准确率,整体平均69.5%,暴露严重短板。
- 适合研究视觉细节理解、模型鲁棒性或评测系统漏洞的团队。
当前最先进的视觉语言模型(VLMs)是否真正具备基础视觉理解能力?我们提出VisualOverload,一个包含2,720个问答对的新型视觉问答(VQA)基准,其真实答案由私有方式保存。与以往侧重全局图像理解的VQA数据集不同,VisualOverload挑战模型在高度密集(即“过载”)场景中完成简单、无需先验知识的视觉任务。数据集基于公共领域的高分辨率画作扫描,画面中包含多个角色、动作及复杂情节,背景细节丰富。我们手动标注了六类任务的问题,以全面检验场景理解能力。我们假设现有基准过度高估了VLM性能,细节编码与推理仍是其难点,尤其在密集场景下。实测显示,37个测试模型中表现最佳的o3模型在最困难测试集上仅达19.6%准确率,整体平均为69.5%。此外,误差分析揭示多种失败模式:计数能力缺失、OCR失效、复杂任务中出现显著逻辑矛盾。VisualOverload揭示了当前视觉模型的关键缺陷,并为社区提供了重要资源以推动更优模型的发展。
原文摘要 · Abstract (English)
Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 question-answer pairs, with privately held ground-truth responses. Unlike prior VQA datasets that typically focus on near global image understanding, VisualOverload challenges models to perform simple, knowledge-free vision tasks in densely populated (or, overloaded) scenes. Our dataset consists of high-resolution scans of public-domain paintings that are populated with multiple figures, actions, and unfolding subplots set against elaborately detailed backdrops. We manually annotated these images with questions across six task categories to probe for a thorough understanding of the scene. We hypothesize that current benchmarks overestimate the performance of VLMs, and encoding and reasoning over details is still a challenging task for them, especially if they are confronted with densely populated scenes. Indeed, we observe that even the best model (o3) out of 37 tested models only achieves 19.6% accuracy on our hardest test split and overall 69.5% accuracy on all questions. Beyond a thorough evaluation, we complement our benchmark with an error analysis that reveals multiple failure modes, including a lack of counting skills, failure in OCR, and striking logical inconsistencies under complex tasks. Altogether, VisualOverload exposes a critical gap in current vision models and offers a crucial resource for the community to develop better models. Benchmark: http://paulgavrikov.github.io/visualoverload
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。