构建首个评估视觉深度搜索能力的基准,测试模型多步推理与反复看图能力。
VistaHop: Benchmarking Long-Horizon Visual DeepSearch

- 设计多轮看图任务,要求模型跨区域追踪细粒度线索。
- 顶尖模型在600个任务上仅达26.33%准确率,表明当前能力严重不足。
- 适合研究视觉推理、智能体系统与多模态交互的学者参考。
视觉深度搜索任务要求多模态大语言模型通过反复检查图像区域,将推理锚定在视觉证据上,并在多个步骤间连接细粒度线索。然而现有基准主要评估单步视觉理解或孤立的视觉问答生成,存在难度低、搜索范围短、仅单次图像浏览等问题,难以评估模型迭代回溯视觉证据和跨步推理的能力。本文提出VistaHop,一个专为评估视觉深度搜索设计的基准,涵盖600张图像、25种视觉搜索场景和600个任务,评估重复图像检查、视觉锚点定位及长跨度证据遍历。我们还提出VistaArena,支持工具化交互(如视觉检索、图像检查、证据推理)的统一评估框架。实验表明,即使最先进的多模态大语言模型仍远未解决该任务,表现最佳的SenseNova-MARS-32B模型仅达到26.33% Pass@1,凸显专用基准与更优智能体方法的重要性。
原文摘要 · Abstract (English)
Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image regions, grounding reasoning in visual evidence, and connecting fine-grained clues across multiple steps. However, existing benchmarks primarily evaluate single-step visual understanding or isolated visual-query response generation. They have limited difficulty, limited search horizons, and single-pass image inspection, and thus fail to evaluate models' ability to iteratively revisit visual evidence and reason across multiple steps. In this work, we introduce VistaHop, a benchmark designed specifically to evaluate Visual DeepSearch. It evaluates repeated image inspection, visual-anchor grounding, and long-horizon evidence traversal across different visual regions. VistaHop comprises 600 images, 25 visual search scenarios, and 600 Visual DeepSearch tasks. We also propose VistaArena, a unified evaluation framework that supports tool-based interactions, including visual retrieval, image inspection, and evidence-grounded reasoning. Experiments show that even state-of-the-art MLRMs remain far from solving VistaHop, with the best-performing model, SenseNova-MARS-32B, achieving only 26.33% Pass@1. These findings highlight the importance of specialized benchmarks and improved agentic methods for Visual DeepSearch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。