arXiv:2606.03273cs.CVcs.AI2026-06

构建首个评估视觉深度搜索能力的基准,测试模型多步推理与反复看图能力。

VistaHop: Benchmarking Long-Horizon Visual DeepSearch

论文配图:VistaHop: Benchmarking Long-Horizon Visual DeepSearch
图 1 · 摘自论文原文
  • 设计多轮看图任务,要求模型跨区域追踪细粒度线索。
  • 顶尖模型在600个任务上仅达26.33%准确率,表明当前能力严重不足。
  • 适合研究视觉推理、智能体系统与多模态交互的学者参考。

视觉深度搜索任务要求多模态大语言模型通过反复检查图像区域,将推理锚定在视觉证据上,并在多个步骤间连接细粒度线索。然而现有基准主要评估单步视觉理解或孤立的视觉问答生成,存在难度低、搜索范围短、仅单次图像浏览等问题,难以评估模型迭代回溯视觉证据和跨步推理的能力。本文提出VistaHop,一个专为评估视觉深度搜索设计的基准,涵盖600张图像、25种视觉搜索场景和600个任务,评估重复图像检查、视觉锚点定位及长跨度证据遍历。我们还提出VistaArena,支持工具化交互(如视觉检索、图像检查、证据推理)的统一评估框架。实验表明,即使最先进的多模态大语言模型仍远未解决该任务,表现最佳的SenseNova-MARS-32B模型仅达到26.33% Pass@1,凸显专用基准与更优智能体方法的重要性。

原文摘要 · Abstract (English)

Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evidence, and connecting fine-grained clues across multiple steps. However, existing benchmarks primarily evaluate single-step visual understanding or isolated visual-query response generation. They have limited difficulty, limited search horizons, and single-pass image inspection, and thus fail to evaluate models' ability to iteratively revisit visual evidence and reason across multiple steps. In this work, we introduce VistaHop, a benchmark designed specifically to evaluate Visual DeepSearch. It evaluates repeated image inspection, visual-anchor grounding, and long-horizon evidence traversal across different visual regions. VistaHop comprises 600 images, 25 visual search scenarios, and 600 Visual DeepSearch tasks. We also propose VistaArena, a unified evaluation framework that supports tool-based interactions, including visual retrieval, image inspection, and evidence-grounded reasoning. Experiments show that even state-of-the-art MLRMs remain far from solving VistaHop, with the best-performing model, SenseNova-MARS-32B, achieving only 26.33% Pass@1. These findings highlight the importance of specialized benchmarks and improved agentic methods for Visual DeepSearch.

视觉推理多模态智能体评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。