首个评估多模态模型‘用图像思考’能力的基准,挑战现实场景下的图像操作与推理。
Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
- 构建动态图像操作任务,让模型像人一样编辑、变换图像来辅助推理。
- 最强模型仅18.68%通过率,说明当前模型在视觉工具融合上严重不足。
- 适合研究多模态智能、具身认知或视觉-语言协同的开发者和研究人员。
多模态大语言模型(MLLMs)正广泛应用于真实场景,但用户提供的图像常不完美,需主动裁剪、编辑或增强以提取关键视觉线索。超越静态视觉感知,MLLMs还需能动态变换视觉内容,并与通用工具结合解决复杂任务。然而,从被动接收视觉信息转向将视觉作为可操作的认知工作区,仍缺乏系统研究。现有评测大多沿用‘思考图像’范式,将图像视为静态输入。为此,我们提出VisualToolBench——首个聚焦‘用图像思考’范式的视觉工具使用推理评测基准,涵盖1,204个开放性、高难度视觉-文本任务(603个单轮,601个多轮),覆盖五个不同领域,每项任务配有详细评分标准,支持系统评估。实验表明,当前主流MLLMs在整合视觉与通用工具方面表现不佳,即使最强模型GPT-5-think也仅达18.68%通过率。我们还观察到工具使用行为差异:OpenAI模型因多样化图像操作显著提升性能,而Gemini-2.5-pro则无改善。VisualToolBench为推动MLLM视觉智能发展提供了关键洞见。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are increasingly applied in real-world scenarios where user-provided images are often imperfect, requiring active image manipulations such as cropping, editing, or enhancement to uncover salient visual cues. Beyond static visual perception, MLLMs must also think with images: dynamically transforming visual content and integrating it with other tools to solve complex tasks. However, this shift from treating vision as passive context to a manipulable cognitive workspace remains underexplored. Most existing benchmarks still follow a think about images paradigm, where images are regarded as static inputs. To address this gap, we introduce VisualToolBench, a visual tool-use reasoning benchmark that rigorously evaluates MLLMs' ability to perceive, transform, and reason across complex visual-textual tasks under the think-with-images paradigm. VisualToolBench comprises 1,204 challenging, open-ended vision tasks (603 single-turn, 601 multi-turn) spanning across five diverse domains, each paired with detailed rubrics to enable systematic evaluation. Our evaluation shows that current MLLMs struggle with tasks requiring effective integration of vision and general-purpose tools. Even the strongest model (GPT-5-think) reaches only 18.68% pass rate. We further observe divergent tool-use behaviors, with OpenAI models benefiting from diverse image manipulations while Gemini-2.5-pro shows no improvement. By introducing the first benchmark centered on think with images, VisualToolBench offers critical insights for advancing visual intelligence in MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。