arXiv:2603.16289cs.CVcs.AI2026-03被引 8

新基准评估网页搜索中视觉信息的推理能力,发现顶级模型准确率仅47.6%。

VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents

  • 构建多阶段人工标注数据集,强化视觉推理评估
  • 真实网页视觉信息参与推理链,提升评估真实性
  • 适合研究多模态搜索、视觉推理的学者与工程师

多模态大语言模型的快速发展使浏览代理能在真实世界中获取并推理多模态信息。但现有基准存在两大缺陷:对视觉推理能力评估不足,且忽略网页原生视觉信息在推理链中的作用。为此,我们提出面向视觉原生搜索的新基准VisBrowse-Bench,包含169个VQA实例,覆盖多个领域,通过文本-图像检索与联合推理实现多模态证据交叉验证,评估模型在搜索过程中的视觉推理能力。数据由人工专家经多阶段流程构建并严格人工校验。我们还设计了一种能有效驱动代理主动收集和推理视觉信息的智能体工作流。在该工作流下,对开源与闭源模型进行综合评估,结果显示,性能最佳的Claude-4.6-Opus模型准确率仅为47.6%,而专有模型o3-deep-research准确率也仅41.1%。代码与数据可于https://github.com/ZhengboZhang/VisBrowse-Bench获取。

原文摘要 · Abstract (English)

The rapid advancement of Multimodal Large Language Models (MLLMs) has enabled browsing agents to acquire and reason over multimodal information in the real world. But existing benchmarks suffer from two limitations: insufficient evaluation of visual reasoning ability and the neglect of native visual information of web pages in the reasoning chains. To address these challenges, we introduce a new benchmark for visual-native search, VisBrowse-Bench. It contains 169 VQA instances covering multiple domains and evaluates the models' visual reasoning capabilities during the search process through multimodal evidence cross-validation via text-image retrieval and joint reasoning. These data were constructed by human experts using a multi-stage pipeline and underwent rigorous manual verification. We additionally propose an agent workflow that can effectively drive the browsing agent to actively collect and reason over visual information during the search process. We comprehensively evaluated both open-source and closed-source models in this workflow. Experimental results show that even the best-performing model, Claude-4.6-Opus only achieves an accuracy of 47.6%, while the proprietary Deep Research model, o3-deep-research only achieves an accuracy of 41.1%. The code and data can be accessed at: https://github.com/ZhengboZhang/VisBrowse-Bench

多模态视觉推理搜索代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。