arXiv:2511.01833cs.CV2025-11被引 29

评测视觉推理模型的图像思维能力,推动智能工具使用发展。

TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning

  • 构建13项任务的综合评估基准,要求模型动态使用工具处理图像。
  • 22个模型测试显示,仅少数具备真正图像思维能力。
  • 适合研究多模态大模型、智能代理与工具调用的学者参考。

视觉推理的前沿正转向如OpenAI o3等模型,它们能智能创建并操作工具以处理图像来解决问题,即链式思考中的“图像思维”。然而现有基准无法全面捕捉这一高级能力。即使是最常见的视觉搜索基准,也仅测试定位、裁剪等基础操作,难以反映复杂、动态且依赖工具的推理过程。我们提出TIR-Bench,一个涵盖13种多样化任务的综合性评估基准,每个任务均需在链式思考中创新使用图像处理工具。评估了22个跨开源与专有、含显式工具增强的多模态大语言模型(MLLMs)。结果表明,TIR-Bench普遍具有挑战性,优秀表现需具备真实的图像思维能力。最后,我们开展初步研究,对比直接微调与代理式微调的效果。

原文摘要 · Abstract (English)

The frontier of visual reasoning is shifting toward models like OpenAI o3, which can intelligently create and operate tools to transform images for problem-solving, also known as thinking-\textit{with}-images in chain-of-thought. Yet existing benchmarks fail to fully capture this advanced capability. Even Visual Search, the most common benchmark for current thinking-\textit{with}-images methods, tests only basic operations such as localization and cropping, offering little insight into more complex, dynamic, and tool-dependent reasoning. We introduce \textbf{TIR-Bench}, a comprehensive benchmark for evaluating agentic thinking-with-images across 13 diverse tasks, each requiring novel tool use for image processing and manipulation in chain-of-thought. We evaluate 22 multimodal large language models (MLLMs), from leading open-sourced and proprietary models to those with explicit tool-use augmentation. Results show that TIR-Bench is universally challenging, and strong performance requires genuine thinking-with-images capabilities. Finally, we present a pilot study comparing direct versus agentic fine-tuning.

视觉推理图像思维多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。