arXiv:2504.18406cs.CL2025-04ICCV被引 3

评测视觉大模型在高分辨率图像理解上的能力,发现普遍表现不佳。

HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?

  • 构建统一基准HRScene,覆盖25类真实与合成高分辨图像。
  • 28个模型平均准确率仅约50%,真实任务表现存在明显差距。
  • 揭示模型难以有效利用图像区域,适合研究高分辨视觉理解者。

高分辨率图像(HRI)理解旨在处理像素数量庞大的图像,如病理图像和农业航拍图,部分图像超过100万像素。尽管视觉大语言模型(VLMs)宣称可处理此类图像,但缺乏全面的评估基准。为此,我们提出HRScene,一个全新的统一基准,涵盖25个真实世界数据集和2个合成诊断数据集,图像分辨率从1,024×1,024到35,503×26,627不等。HRScene由10名研究生级标注员收集并重新标注,覆盖微观至放射影像、街景、远距离图像及望远镜图像等25种场景,包含真实物体、扫描文档及多图像合成内容。两个诊断数据集通过组合目标图像、正确答案与干扰图像的不同顺序生成,用于评估模型对高分辨率图像区域的利用能力。我们对28个VLM进行了实验,包括Gemini 2.0 Flash和GPT-4o。结果显示,当前VLM在真实任务上平均准确率约为50%,暴露出显著的能力差距;合成数据实验表明,模型在区域利用上存在严重偏差与‘中间迷失’现象,为未来研究提供重要方向。

原文摘要 · Abstract (English)

High-resolution image (HRI) understanding aims to process images with a large number of pixels, such as pathological images and agricultural aerial images, both of which can exceed 1 million pixels. Vision Large Language Models (VLMs) can allegedly handle HRIs, however, there is a lack of a comprehensive benchmark for VLMs to evaluate HRI understanding. To address this gap, we introduce HRScene, a novel unified benchmark for HRI understanding with rich scenes. HRScene incorporates 25 real-world datasets and 2 synthetic diagnostic datasets with resolutions ranging from 1,024 $\times$ 1,024 to 35,503 $\times$ 26,627. HRScene is collected and re-annotated by 10 graduate-level annotators, covering 25 scenarios, ranging from microscopic to radiology images, street views, long-range pictures, and telescope images. It includes HRIs of real-world objects, scanned documents, and composite multi-image. The two diagnostic evaluation datasets are synthesized by combining the target image with the gold answer and distracting images in different orders, assessing how well models utilize regions in HRI. We conduct extensive experiments involving 28 VLMs, including Gemini 2.0 Flash and GPT-4o. Experiments on HRScene show that current VLMs achieve an average accuracy of around 50% on real-world tasks, revealing significant gaps in HRI understanding. Results on synthetic datasets reveal that VLMs struggle to effectively utilize HRI regions, showing significant Regional Divergence and lost-in-middle, shedding light on future research.

高分辨率图像视觉大模型评测基准区域利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。