测试大模型在密集场景中是否真能主动寻找关键视觉细节。
VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes

- 设计高密度场景任务,迫使模型定位微小视觉线索。
- 工具启用下模型最高仅达56.01%准确率,远低于人类63.00%。
- 通过黑块替换中间图像验证模型依赖真实视觉证据。
前沿多模态大语言模型在细粒度感知基准上报告超过90%的准确率,但这些分数未必反映对视觉证据的真实使用。已有研究指出三大性能虚高陷阱:语言先验和问题中的词汇线索使模型无需看图即可推断答案;视觉编码器的粗粒度全局语义可绕过精细局部细节;部分“带图思考”基准中,篡改工具返回的中间图像几乎不影响最终答案。这表明单纯提升输入分辨率或扩大问题池无法激发真正的主动视觉搜索。为此,我们提出VisualNeedle,一个高信息密度、细粒度的基准,其关键证据被限制在极小空间区域,难以一眼察觉。我们还引入反事实裁剪-黑化设置,用同尺寸黑色图像替换工具返回的裁剪区域,检验工具性能是否真正依赖中间视觉证据。我们在三个设置下评估9个主流MLLMs:无工具、标准工具启用、裁剪-黑化。无工具准确率始终低于20%,最佳工具模型仅达56.01%,仍低于人类多数投票的63.00%。结果揭示模型在细粒度视觉搜索中的持续局限性,而黑化消融实验确认VisualNeedle上的成功确实依赖于真实的中间视觉证据。
原文摘要 · Abstract (English)
Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such scores do not necessarily imply faithful use of visual evidence. Prior studies have identified three shortcuts that inflate benchmark performance. First, linguistic priors and lexical cues in questions often enable models to infer plausible answers without seeing the image. Second, coarse global semantics from the visual encoder can bypass fine-grained local details. Third, in some ``think-with-images'' benchmarks, corrupting the intermediate images returned by visual tools barely affects the final answer. These findings suggest that higher input resolution or larger question pools alone do not elicit genuine active visual search. To address this, we introduce VisualNeedle, a challenging, information-dense, and fine-grained benchmark for scenes where critical evidence is spatially constrained to minute regions and not discernible at a glance. We further propose a counterfactual crop-black setting, which replaces crops returned by tools with black images of the same size, to test whether tool-enabled performance truly relies on intermediate visual evidence. We evaluate 9 promninent MLLMs across three settings: no-tool, standard tool-enabled, and crop-black. No-tool accuracy stays below 20\%, and the best tool-enabled model reaches only 56.01\%, still trailing the 63.00% human majority-vote accuracy. These results reveal persistent limitations in fine-grained visual search, while the crop-black ablation confirms that success on VisualNeedle hinges on genuine intermediate visual evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。